I created an interactive digital avatar of myself — and you can talk to it
When TechCrunch’s reporter got a hands‑on demo of Synthesia’s newest offering—a fully interactive digital avatar of himself—the experience felt like a preview of a future where your own virtual twin could field questions, deliver presentations, and maybe even take a corner office.
When TechCrunch’s reporter got a hands‑on demo of Synthesia’s newest offering—a fully interactive digital avatar of himself—the experience felt like a preview of a future where your own virtual twin could field questions, deliver presentations, and maybe even take a corner office. The demo, staged at Synthesia’s freshly opened New York studio, turned a routine PR pitch into a live‑action test of how far AI‑driven avatars have come, and what that could mean for journalists, marketers and the broader workforce.
From a static video to a talking twin
Synthesia, the U.K.–based video‑generation startup that recently hit a $4 billion valuation, has been known for letting enterprises create polished video messages with lifelike avatars. Its core platform lets users type a script, pick an avatar, and export a video that looks as if a human presenter is speaking. The company’s latest push, however, adds a conversational layer: the “Roleplay Sessions” product that lets employees practice sales pitches or interview scenarios with an avatar that can listen, respond and even score the interaction.
During the September office opening, the reporter was invited to become the first journalist outside of Synthesia’s own head of corporate affairs, Alexandru Voica, to receive a personal avatar. The process was straightforward but high‑tech: a mini‑studio captured dozens of photos and a two‑minute voice recording, then the team built both a “personal” avatar that simply reads any supplied script and an “interactive” avatar that can understand spoken queries and reply using a deterministic language model trained on a single article about venture‑backed startup fraud.
How the avatar actually works
The interactive twin runs on a stack that stitches together several AI components. First, a voice‑to‑text model transcribes what a user says. Next, an agentic language model parses the text, decides on an appropriate action, and generates a response limited to the pre‑trained content. That response is then fed to a text‑to‑voice engine, which can be one of Synthesia’s own voice models or an alternative from providers such as Cartesia, ElevenLabs, Google or OpenAI. Finally, Synthesia’s proprietary video model animates the avatar’s facial movements to match the audio output.
Customers can host the resulting avatars on any cloud platform they prefer, or pay Synthesia to manage hosting. This flexibility mirrors the company’s broader API offering, which lets developers pull together Synthesia’s video and voice models with other services to build custom interactive experiences.
The demo: a mix of novelty and eeriness
When the reporter first tried the personal avatar, he typed a simple script about New York’s autumn and watched the digital double recite it with a voice that sounded “fairly accurate” and even avoided the hoarseness present in his original recording. Friends who saw the playback described the result as “interesting and creepy,” noting that the visual likeness was close enough to be unsettling but not a perfect replica.
The interactive version, however, felt more like a controlled chatbot. Because it was trained only on the venture‑fraud article, every question—whether about the reporter’s previous employer or his Manhattan neighborhood—was met with a polite redirection back to that story. Even his parents, who tried to stump the avatar with personal trivia, received the same scripted answer. The deterministic nature of the model meant it never strayed beyond its training data, a feature that both reassured and unnerved the onlookers.
Why deterministic matters (and why it doesn’t solve everything)
Determinism keeps the avatar from hallucinating or providing inaccurate information, a crucial safeguard for corporate use cases where brand reputation is on the line. Synthesia’s own description of the three product tiers—classic scripted videos, the agentic “Sessions” platform, and the API for custom builds—highlights this safety net: the interactive avatars are only as knowledgeable as the content they’re fed.
But the reporter points out a lingering gap: trust. In journalism, credibility hinges on the reporter’s judgment, sources and personal nuance—qualities that a deterministic avatar can mimic only superficially. An investor quoted in the conversation dismissed the idea of AI‑driven news anchors outright, while others remained ambivalent, suggesting that avatars could augment but not replace human storytellers.
Business implications for Synthesia and its rivals
Synthesia’s move into interactive avatars puts it in direct competition with other digital‑twin players such as D‑ID, HeyGen and Colossyan, all of which have been racing to add conversational capabilities to their portfolios. By offering a modular stack that lets enterprises plug in third‑party voice models and host on any cloud, Synthesia positions itself as a flexible, enterprise‑grade option—a strategy that aligns with its reported $100 million annual recurring revenue milestone reached last year.
The company’s ability to spin up a custom avatar in “a couple of days” also signals a scalable production pipeline. If corporate clients can quickly generate branded twins for internal training, sales enablement or customer support, the revenue upside could be significant, especially as more firms look to automate repetitive communication tasks without sacrificing a human‑like presence.
Potential societal ripple effects
Beyond the boardroom, the technology raises broader cultural questions. The reporter muses that a personal avatar could “answer questions” while you’re on vacation, effectively extending your professional persona around the clock. That convenience is tempting, yet it also blurs the line between authentic self‑presentation and algorithmic performance.
There’s also a psychological angle. The demo’s “creepy” factor hints at a possible “AI psychosis” where users begin to attribute agency or emotions to deterministic avatars that never truly understand. The reporter notes that a nondeterministic version—one powered by an open‑ended chatbot—could amplify that effect, letting the twin pontificate beyond its training and potentially erode the user’s sense of reality.
What the future might hold for digital twins
As Synthesia rolls out its interactive avatars to more enterprise customers, the technology will likely seep into everyday workflows: sales reps rehearsing pitches with a virtual counterpart, HR teams onboarding new hires via a friendly digital guide, and perhaps even newsrooms experimenting with AI‑assisted reporting. The reporter’s own mixed feelings capture the ambivalence many feel—fascination with the novelty, but caution about the deeper implications.
What remains clear is that the barrier to creating a believable digital twin is dropping fast. Within days, a startup can capture a few minutes of audio, a handful of photos, and spin up a talking version of a person that can field scripted queries with a polished voice. Whether that will translate into widespread adoption—or into a backlash over authenticity—will depend on how companies, regulators and the public negotiate the trade‑off between efficiency and trust.
This article was produced with AI-assisted research and editorial support. Reporting is based on the source material cited below. Sources: TechCrunch; techcrunch.com; Global1.News (27 September 2026).
By Nova Chen, Staff Writer
What's Your Reaction?
Like
0
Dislike
0
Love
0
Funny
0
Wow
0
Sad
0
Angry
0
Comments (0)