cara-4, and the product around the model
Today Anam is releasing cara-4, our latest real-time avatar model.
When I wrote about cara-3, I focused on the model: realism, latency and the work required to run generated video in a live conversation. With cara-4, I keep coming back to the product around it. A model can be much better and still feel like a small update if people can't see or use what changed.
The research team built a model with a much wider emotional range, richer facial expressions and support for portrait video. My part was helping those capabilities make sense in the product.
A face with more range
cara-3 made our avatars feel considerably more natural. cara-4 gives them more range.
The difference is easiest to spot when the emotional tone changes. A cheerful explanation shouldn't look like a difficult apology, and a calm tutor should move differently from an excited sales agent. cara-4 can carry more of that intent through the eyes, brows, mouth and small movements of the head, without losing the identity of the person on screen.
That matters because a voice agent already has a tone. Its prompt, voice and choice of words create a character, whether the developer planned one or not. The face needs to belong to the same character.
Photorealism is only one option
cara-4 also works with stylised characters. In our new Digital Clone flow, one photo can stay realistic or be restyled as an Animated 3D character before it becomes a live avatar. There's no separate rigging or animation step.
I tried it on myself. Seeing the 3D version of my face move and answer back is one of the stranger demos I've made at Anam. You can talk to him here:
Giving developers a way to direct it
More expressive output is useful, but it also creates a product question: how should a developer ask for the performance they want?
Our answer started internally as Director Notes. In the product, it appears as style and expressivity controls. You can choose the emotional style for a persona, set how strongly it should come through, and use inline cues when the performance needs to change during a response.
The controls look simple. Getting there wasn't. The team had to connect model behaviour, API design and sensible defaults so that increasing expressivity feels predictable rather than random. The result is a face that can follow the direction of an experience, instead of every avatar delivering every line in roughly the same register.
I like this feature because it gives product teams room to make choices. A children's tutor, a support agent and a character in a game should not all perform in the same way.
Portrait mode changes more than the frame
Portrait output sounds like a width-and-height change. It touches almost everything around the video.
I worked on much of the Lab UX for it: the orientation control, checkpoint-aware availability, full-screen behaviour, square source capture and previews that show the exact landscape and portrait crops the engine will produce. We added resolution guidance too, because discovering that a face has been framed badly after creating the avatar is an expensive way to learn about aspect ratios.
One of the less glamorous bugs was a portrait toggle that could appear for a model checkpoint that could not render portrait video. The interface offered a choice, then the session failed. Fixing it meant teaching the UI about the checkpoint the session would actually use, and making the server return a useful error if an impossible request still got through.
That work is invisible when it's right. You choose portrait, see an honest preview and start the call.
Rebuilding the route into the product
I have also spent the past week rebuilding the main Personas experience in Lab. The new home is organised around exploring personas, talking to one immediately, and then adding or creating one of your own. The call experience is video-first and designed around portrait footage, including on a phone.
The same work adds two creation routes. Persona Design can start from a written description. Digital Clone guides someone through a photo or webcam capture, voice, knowledge and expressivity before opening a live call. There are plenty of unromantic details behind those screens: camera permissions, session ownership, loading states, mobile transcript layout and cleaning up microphones when a dialog closes. They are also the difference between a capability and a usable product.
A user never experiences a model checkpoint. They experience the crop, the defaults, the controls and whether the call starts.
I used to think a sufficiently better model would announce itself. It doesn't. People encounter it through a series of product decisions, most of them small. The job is to make the new capability obvious without making the machinery complicated.
Try cara-4
You can read Anam's cara-4 announcement for the company view and a live example.
Or sign up and build with cara-4. Try the same persona in landscape and portrait, move the expressivity control, and give it a style that actually fits the job. That's where the change becomes clear.