CASE_ID: MAIA

The MAIA Experience

My NYU Tandon master’s capstone explored how people experience privacy and trust with embodied AI. I designed and built a locally hosted character that sees, listens, and converses in a physical story world, and evaluated the experience with 37 participants. The project connects AI engineering with interaction design and live experience.

ROLELead Designer & Engineer
DOMAINPrivate Embodied AI · Research & Design Engineering
STACK
GenAIVoice UIPythonLLM
DEPLOYMENTthe-maia-experience.framer.ai
The Challenge

"A ten-second wait disrupted the human quality of the encounter. I needed to improve response time while preserving the pacing, privacy, and trust of a live conversation with an embodied AI character."

My responsibility

NYU Tandon master’s capstone. Lead designer and engineer for the embodied AI experience: experience design, prompt engineering, the multilayered real-time pipeline, and the operational interface. Evaluated the experience with 37 participants.

Judgment & tradeoffs

Decisions and why

  1. 01

    Replace a sequential wait with overlapping work

    The first version took around ten seconds to respond, which did not support the human quality of the encounter. I engineered a multilayered pipeline so inference could run while other scripted elements were running. Response times came down to approximately 200 milliseconds to three seconds, depending on the interaction.

  2. 02

    Treat privacy and trust as experience decisions

    I chose local processing and designed listening and thinking cues into the physical experience. The goal was for participants to understand the character’s behavior and retain agency while I evaluated usability, privacy, and trust with 37 participants.

System Architecture

How it Works

A locally hosted multimodal stack combining computer vision, speech recognition, a language model, and voice synthesis. I engineered a multilayered pipeline that allowed inference to run while other scripted elements were running, bringing response times from around ten seconds to approximately 200 milliseconds to three seconds. A custom Python backend coordinated the experience, and an operational interface gave me control of the AI, audio, vision, and lighting.

Key Components
  • 1Immersive production design that established the story world and the character’s believability.
  • 2Dialogue states changed with audio and visual triggers. (camera → STT → LLM → TTS → led lighting).
  • 3Overlapping inference and scripted elements to support conversational pacing.
  • 4Privacy controls that prevented responses from being automatically saved to the cloud.
Design & Implementation

Shaping the Experience

Really there were two interfaces: 1. The operation of the audio/visual/AI experience. 2. The live audio chat interface with the AI. To mitigate the 'Uncanny Valley,' I designed the interface to be an invisible layer. Instead of a chat window, the UX relied on LED cues—lighting changes and subtle sound design—to signal the AI's 'listening' and 'thinking' states. This reduced the cognitive load of a standard conversational UI, allowing users to maintain eye contact with the physical avatar. The second interface (shown right) is the workspace I custom made to give me operational control of lights, computer vision inputs and outputs, inference, and audio (DAW). The left image is the prototype wireframe and the right is the final product UI.

Outcomes & Impact
  • 01

    Reduced response time from around ten seconds to approximately 200 milliseconds to three seconds by overlapping inference with scripted elements in a multilayered pipeline.

  • 02

    Designed 'thinking state' animations that maintained narrative immersion during processing.

  • 03

    Personalized interaction pacing and conversation based on discussion with AI.

  • 04

    Processed voice data locally to support privacy in the live AI experience.

  • 05

    Produced personalized encounters and AI-generated physical mementos based on visitor input.