NVIDIA Spark Hackathon Seattle 2026 · Winner

MEMO.

A visual memory assistant that remembers where things are, answers everyday questions, and brings a trusted person to help when it matters.

Local-first and private by default. Built for our parents, on a single Acer GN100.

Not a cloud model answer. Reasoning, perception, speech, and memory run on the device.

Built with NVIDIA · Runs on Acer

“Where did I leave my keys?”

“On the living-room coffee table at 10:42, based on the last confirmed placement. I have not confirmed a newer location.

The problem

Being usefully uncertain matters more than sounding confident.

We build MEMO for our parents.

It helps them keep track of important belongings, such as keys or a wallet, so they can ask where something is and get its last confirmed location. It gives them a simple, hands-free way to ask everyday questions.

But not every situation should be left to AI. When judgment, reassurance, or human reasoning is especially valuable, MEMO connects our parents with family members they know and trust. It does not try to replace family support. It keeps that support within reach.

Our goal is to give our parents greater confidence and independence while giving their families peace of mind, and a trusted way to help when it matters.

What it does

Three ways MEMO helps, on one local backend.

Each experience is a real actor journey. The system does not just answer, it knows what it does not know.

Remember

Visual memory

A parent asks where an object is. MEMO returns the last confirmed placement, honest about whether that location is current, historical, ambiguous, or unknown.

Outcome: confidence in everyday life, not a guess presented as fact.

Answer

Local assistant

A parent asks out loud. MEMO listens, reasons on device, and replies privately. No account, no cloud, no screen required.

Outcome: easy access to answers, on their terms, in their home.

Connect

Remote human help

A parent says “call my remote assistant.” A trusted family member’s phone rings. They join to see what the parent sees and speak with them.

Outcome: human judgment within reach, exactly when it matters.

See it in action

Watch the demo.

A real run on the glasses: register an object, ask where it is, then bring a remote helper in. It plays right here.

The stack

A single local NVIDIA stack.

Every capability runs on one device. Private by default, no hosted model required. Each model links to its page.

Agent reasoning and tool routing
Multimodal temporal reasoning
Personal-object recognition
Text-to-speech
Glasses media
RayNeo X3 Pro via LiveKit / WebRTC
Trusted object state
Application Memory (structured observations and evidence)

Where it stands

Built and working. Honest about what is not.

Open an item to see the technical detail behind it.

Working today

Glasses app
RayNeo client for pairing, camera and microphone publishing, HUD events, and return audio.
Remote helper app
Paired mobile client for incoming assist requests and LiveKit human calls.
Operator Console
Live video, memory review and reset, enrollment, speech, agent, and assist UI.
Media Gateway
LiveKit session authority, bounded media relay, and remote-assist lifecycle. Retains no raw media.
Application Memory
Durable object registry, evidence and state reducer, cross-session lookup.
Agent, Speech, Vision
Local Nemotron orchestration, Parakeet STT and Kokoro TTS, Cosmos and C-RADIO recognition.

On the roadmap

Fully automatic enrollment
Operator-guided today; automatic crop extraction is a measured next step.
Continuous visual tracking
The system works on sparse event windows, not live motion state.
Barge-in
The wearer waits for the current answer before starting the next turn.
Object identity beyond the gallery
Recognition is bounded by the enrolled set; no match means no memory write.

The projected Console is the proof surface. MEMO is a memory aid, not a guarantee of an object’s current location, and not a safety-critical system.

How it runs

Built to keep everything local.

Wearer
RayNeo X3 Pro
Session authority
Media Gateway
Local inference
Perception · Speech · Reasoning
Trusted state
Application Memory
Perception
  • Cosmos 3 Nano on sparse windows
  • C-RADIOv4-H object recognition
  • Operator-confirmed crops become trusted placements
Speech
  • Parakeet TDT transcribes
  • Kokoro speaks the reply
  • One bounded, local voice pipeline
Reasoning and memory
  • Nemotron 3.5 Lightning routes tools
  • Application Memory is authoritative about location
  • Answers are evidence-bounded

When help is needed, the wearer’s request opens a LiveKit room to a paired helper phone. Human judgment stays within reach.

Privacy by design

Your video stays with you. We keep only evidence.

MEMO never stores full-length video in the backend. The media gateway retains no raw media, and perception keeps only a short, bounded window of frames it has already sampled and relayed, evicting anything older than that window. What survives is evidence.

No full-length video

The media gateway is built with a hard zero for raw media. The live stream is never recorded to the backend.

Short event evidence

When the camera confirms an object such as a keychain, the system keeps a still frame and a short clip of that event window, a few seconds long, as the evidence of what it saw.

Proof when it matters

That evidence is stored with the placement in the local memory service. Ask later, “where is my keychain?” and MEMO shows the last confirmed location with the evidence to back it up when a text answer alone is not enough.

Everything else stays local: the raw stream, transcripts, and inference never leave the device. Keeping a minimal evidence clip is the deliberate exception that lets the memory answer honestly.

The team

The people building MEMO.

Reach any of us directly to learn more.

Alexander Kuznetsov
Nadine Chernova
Erin Shih
Jacky Huang

Message any of us directly about what we are building.