Back to projects

AI / Robotics

Looking Alive

An interactive, perceptive Leonardo da Vinci.

A character that doesn't just talk at you. It notices you, and has the restraint to know when not to.

Active Research~26 FPSSingle GPU
Computer VisionGaze EstimationMulti-Person TrackingRestraint PolicyExplainable AIReal-Time

Every decision on the record: who it engaged, who it deliberately passed over, and which signals drove each call. Nothing is a black box, “why not her?” has an answer.

…strong model implementations of the deep neural network and GMM-HMM, with clear optimizations.

Kimberley Merritt · Academic Award of Excellence

The build, in five chapters

One character across four bodies, plus the idea that ties them together. Two chapters have work you can read today: The Mind is built and running, and The Face is underway. The last two are mapped. Open any chapter below.

// chapter 00 · the idea

The Vision

The throughline
Engagement feels alive when it is earned and intentional, not constant. A character that notices, engages with restraint, and can prove why, reads as present — not as a greeter.
// concept: the four bodies, one characterconcept art

One character, four bodies: screen → face → simulation → robot.

// where the idea starts

Walk up to most digital characters and they run a script. The naive “reactive” version is worse in a real space: it greets whoever is nearest, which is a nuisance. The bet behind this project is narrower, and I think truer: a character feels alive when it does the harder, more human thing, it chooses — reading who is genuinely present with it, engaging the one person for whom it is actually the right moment, and, just as often, choosing to wait. That single shift, from broadcasting to earned, intentional engagement, is the whole idea. It's built for the place it matters most: guests in a real space, a queue, an exhibit, a lobby, where engaging the right person and leaving everyone else in peace is the difference between presence and a nuisance.

// why Leonardo

The persona is Leonardo da Vinci on purpose. He was history's greatest observer, with notebooks full of how light falls, how the eye reads depth, how a face moves. So the character's superpower and his identity are the same thing: noticing, and having the restraint to know when not to. What he says is drawn from his own notebooks and the historical record, and it is always about the shared craft, how a painted gaze seems to follow you across a room, why he left so much unfinished, never a remark about the person in front of him.

// how it comes together

One character, built in four chapters. Each keeps the same beating heart (perceive, decide with restraint, react believably) and changes only the body it lives in: from a screen, to a physical face, through simulation, into a robot that shares the room with you. The engine is the product; Leonardo is the first host.

The same heart, end to endperceive → decide with restraint → react believably

Two chapters have work in them today. The perception and restraint engine, the part that actually does the judging, is live and auditable in The Mind — and the face it drives is rendering now in The Face.

// the research behind every chapter

The PhD arc is built to bring four researchers' strengths together: a believable, deployable character that perceives and reasons about people, explainably, then steps off the screen into a physical, reactive robot.

Markus Gross

ETH Zürich / Disney Research

Interactive digital characters and the technology that makes them feel present, including projection into physical space.

Joseph Campbell

Purdue, CAMP Lab

Theory of mind, anticipating human intent, and interpretable interaction, the backbone of the “explain every decision” principle.

Heni Ben Amor

Arizona State, Interactive Robotics Lab

Reactive control and robot learning: characters and robots that respond to people in the moment, the engine behind the robotic phase.

Stelian Coros

ETH Zürich, Computational Robotics Lab

Physics-based, expressive character and robot motion, how a believable performance transfers to a body that obeys physics.

// why it matters

The single most repeatable bit of theme-park magic is a character who makes a guest feel seen. Today that depends on a gifted human performer. This builds it as a real-time, repeatable, explainable system: a character that notices the specific guest in front of it, reacts in persona, plays to a crowd, and eventually steps off the screen into the room. Da Vinci is the first host; the perception and decision engine is the product.

// selected references

A curated selection from a maintained annotated bibliography of 60+ sources, the research grounding plus the third-party methods the build stands on.

Research grounding

  1. [1]Wampfler, R., et al. (2025). A Platform for Interactive AI Character Experiences (Digital Einstein). SIGGRAPH Conf. Papers '25.
  2. [2]Campbell, J. & Ben Amor, H. (2017). Bayesian Interaction Primitives: A SLAM Approach to Human-Robot Interaction. CoRL, PMLR 78.
  3. [3]Campbell, J., Stepputtis, S. & Ben Amor, H. (2019). Probabilistic Multimodal Modeling for Human-Robot Interaction Tasks. RSS. arXiv:1908.04955.
  4. [4]Oguntola, I., Campbell, J., Stepputtis, S. & Sycara, K. (2023). Theory of Mind as Intrinsic Motivation for Multi-Agent RL. ICML Workshop. arXiv:2307.01158.
  5. [5]Zhang, X.-J., et al. (2025). Model-Agnostic Policy Explanations with Large Language Models. COLM. arXiv:2504.05625.
  6. [6]Serifi, A., et al. (2024). Robot Motion Diffusion Model (RobotMDM): Motion Generation for Robotic Characters. SIGGRAPH Asia.
  7. [7]Coros, S., et al. (2013). Computational Design of Mechanical Characters. ACM TOG 32(4), SIGGRAPH.
  8. [8]Bates, J. (1994). The Role of Emotion in Believable Agents. Communications of the ACM 37(7).

Methods & systems

  1. [9]Cheng, T., Song, L., Ge, Y., et al. (2024). YOLO-World: Real-Time Open-Vocabulary Object Detection. CVPR. arXiv:2401.17270.
  2. [10]Jocher, G., et al. (2024). Ultralytics YOLO11 (software).
  3. [11]Zhang, Y., Sun, P., Jiang, Y., et al. (2022). ByteTrack: Multi-Object Tracking by Associating Every Detection Box. ECCV. arXiv:2110.06864.
  4. [12]Abdelrahman, A. A., et al. (2022). L2CS-Net: Fine-Grained Gaze Estimation.
  5. [13]Lin, T.-Y., Maire, M., Belongie, S., et al. (2014). Microsoft COCO: Common Objects in Context. ECCV.
  6. [14]Glas, D. F., Shiomi, M., Kanda, T., et al. (2017). Personal Greetings: Personalizing Robot Utterances Based on Novelty of Observed Behavior. Int. J. of Social Robotics.
EOF

Joey Schnepel · Phoenix, AZ · 2026