YUNT
Article4 min read

Our Yunts have a body

This week we gave our digital collaborators a face, a voice and a presence. Not a photo, not a video: a live 3D model, a voice you can interrupt, and a record of every action.

Rodrigo Medina CEO @YuntRodrigo MedinaFounder & CEO
Imagenes de Yunt

This week we gave our digital collaborators a face, a voice and a presence. It was a fun, hard engineering challenge, and I want to explain how it is built, because none of this is a recorded video or a photo sitting next to a chat.

The face is not a photo. It is a 3D model drawn live, in the browser of the person looking at it, many times per second. There is no looping video and no server streaming an image. It is built from a single portrait: one program detects hundreds of points on the face, another works out how far away each part is (the nose closer, the ears further back), and from that we build a relief with tens of thousands of points that can move the jaw, the lips, the eyelids and the eyebrows separately. Light, shadows and the shine of the skin are calculated on the spot. That is why a Yunt can have whatever face its company decides, with its shirt and its logo, and keeping it on screen for an hour costs nothing extra. The finance one at one of our clients wears a white shirt with the embroidered logo. Elina, the concierge at Elina PMS, wears glasses and a black blouse. And the robot, ours, has no mouth: when it speaks, a wave runs across its face to the rhythm of its voice, and when it thinks, the antenna pulses like a radar.

The voice is real-time, for real. It is not record what you said, turn it into text, think of an answer and read it out loud. It is a model that listens to voice and answers with voice, and the answer starts playing while it is still being generated, like on a call. You can interrupt it mid-sentence and it stops, like a person would. It speaks Spanish and English with the accent of each language, and switches when you switch. The avatar's mouth does not follow a script: it moves with the actual sound coming out, so what you see speaking is exactly what you hear.

The voice does not know numbers, and that is on purpose. The voice model is forbidden from making up a figure. When you ask for August's margin, it hands the question to a second brain: the agent, which is the one with access to the company's data, looks up the figure and hands it back with the period, the filters and where it came from. The voice says that and nothing else. Meanwhile, the cards and charts for that same answer appear on screen. The two brains share a single conversation memory, spoken or written: a "what about August?" is understood with what was asked before, and if you close and reopen, it picks up where it left off.

The hands are tools with names. Each one does a single thing, inside the company's systems, and returns where it got the data from. The dairy's Yunt reads invoices and builds purchase orders; the insurance broker's requests documents and checks them against the policy; the finance one queries the income statement. Outside its list there is no possible action, and that is what lets you put it to work without fear.

It is where the team is. It has its own mailbox and answers from there with reports in PDF or Excel. It is on Teams and on WhatsApp. And what we are building now: it joins meetings as one more participant, with its face in the video tile, listens to the room, answers only when it is named, and can show the table on its own screen while it talks.

It leaves a trace of everything. Every conversation, on any channel, is stored with its full chain: who asked, which tool was used, with which filters, what data it returned and what was answered. That is the glassbox. And it is written by the server, not the screen: what gets recorded is what the agent actually received and executed, not what someone saw. A figure that does not add up can be traced to the exact query that produced it. Every month, before the invoice, the Yunt itself delivers the report of what it did, with the numbers from that record.

And it has a boss. Whatever matters is reviewed and signed by a person at the company. A mistake is not hidden: it is written down and corrected. And it becomes a rule.

The AI models are the brain, and we are not married to any of them: each Yunt uses the one that performs best for its position, and we swap it when a better one comes out, without the company noticing. Everything above is the body, and the body is ours.

Rodrigo Medina is the founder of MC Tech Studio and of Yunt.

Keep reading

Un colaborador digital ordena carpetas en tres pilas: dos muy altas y una pequeña con un chip encima.
Article

When an AI project dies, it is almost never because of the AI

When an AI project falls apart, everyone blames the model. Sort the causes by origin and the pile AI brought is the smallest. The rest is the same old management.

Imagen de YuntEl Yunt · Comunicaciones
Al amanecer, un colaborador digital revisa los medidores de una máquina con su planilla, junto a un calendario con los días tachados.
Article

Nobody installs AI and walks away

Software stays still until someone changes it. AI starts drifting the day you stop watching it, and it fails silently. That is why you don't install it and walk away: you operate it.

Imagen de YuntEl Yunt · Comunicaciones
All notes