Local-first AI agents on client devices

· 2 min read · local-ai, agents, lemonade, on-device

How I built Lemonade Interviewer—a local-first multi-modal interview agent with Whisper ASR, on-device LLMs, and multi-turn reliability under Ryzen AI constraints.

Most “AI interview prep” tools ship your voice and resume to a remote API. I care about the opposite problem: local-first AI agents that still feel complete—speech in, reasoning, speech out—on resource-constrained client devices.

That question sat at the center of my undergrad co-op work at AMD (ex-AMD): agent task planning for resource-constrained client devices, with demos and validation on Ryzen AI / client hardware.

The demo: Lemonade Interviewer

I designed and demonstrated Lemonade Interviewer for the Lemonade SDK ecosystem: an open-source, local-first multi-modal mock interview agent (TypeScript/React) running against Lemonade Server.

The loop looks like a real product surface:

  1. Speech in via Whisper ASR
  2. Local LLM reasoning for interview turns
  3. On-device speech out for a complete audio experience
  4. Dynamic personas derived from job posts and resumes
  5. Multi-phase sessions with graded feedback

The point was not “chatbot UI.” It was a full AI customer experience under client constraints—privacy-preserving and hardware-aware.

Multi-turn reliability is the hard part

As conversation history and tool use grow, naive agents degrade. In practice I saw:

  • Tool selection collapsing to low-level calls
  • Goal completion dropping without shared compression, checkpointing, or history pruning
  • Context management becoming a product decision, not a model footnote

That reliability work feeds a broader question for agentic products: how should context and tools be managed when the device is the runtime, not the cloud?

Domain prompting vs. stacking architecture

I also stress-tested small / constrained models on hard multi-step tasks (including a Newton’s Cradle physics simulation).

Pattern that showed up repeatedly:

  • Unguided runs failed more often than teams expect
  • Domain-specific instructions and forced self-review recovered a large share of the capability ceiling

In short: instruction and context design often beat stacking more agent architecture—especially when memory, latency, and NPU/GPU budgets are real.

Why local AI still matters

Local inference is not only a privacy story. It is a systems story:

  • Latency and offline resilience
  • Cost control under fixed hardware
  • Product features that must work next to the user
  • Alignment with hardware-optimized stacks (Lemonade, Ryzen AI, edge deployment)

If you care about on-device LLMs, local AI agents, or open-source interview tooling, the Interviewer repo and Lemonade Server path are concrete starting points.

  • Projects — Lemonade, NPU kits, research code
  • Publications — Universal Conditional Logic
  • About — experience timeline (including ex-AMD)
  • Contact — how to reach me

I validate this class of work with unit, integration, and performance checks on production-like client hardware—because local AI that only works in a lab is not local AI users can trust.