Local-first AI agents on client devices
How I built Lemonade Interviewer—a local-first multi-modal interview agent with Whisper ASR, on-device LLMs, and multi-turn reliability under Ryzen AI constraints.
Most “AI interview prep” tools ship your voice and resume to a remote API. I care about the opposite problem: local-first AI agents that still feel complete—speech in, reasoning, speech out—on resource-constrained client devices.
That question sat at the center of my undergrad co-op work at AMD (ex-AMD): agent task planning for resource-constrained client devices, with demos and validation on Ryzen AI / client hardware.
The demo: Lemonade Interviewer
I designed and demonstrated Lemonade Interviewer for the Lemonade SDK ecosystem: an open-source, local-first multi-modal mock interview agent (TypeScript/React) running against Lemonade Server.
The loop looks like a real product surface:
- Speech in via Whisper ASR
- Local LLM reasoning for interview turns
- On-device speech out for a complete audio experience
- Dynamic personas derived from job posts and resumes
- Multi-phase sessions with graded feedback
The point was not “chatbot UI.” It was a full AI customer experience under client constraints—privacy-preserving and hardware-aware.
Multi-turn reliability is the hard part
As conversation history and tool use grow, naive agents degrade. In practice I saw:
- Tool selection collapsing to low-level calls
- Goal completion dropping without shared compression, checkpointing, or history pruning
- Context management becoming a product decision, not a model footnote
That reliability work feeds a broader question for agentic products: how should context and tools be managed when the device is the runtime, not the cloud?
Domain prompting vs. stacking architecture
I also stress-tested small / constrained models on hard multi-step tasks (including a Newton’s Cradle physics simulation).
Pattern that showed up repeatedly:
- Unguided runs failed more often than teams expect
- Domain-specific instructions and forced self-review recovered a large share of the capability ceiling
In short: instruction and context design often beat stacking more agent architecture—especially when memory, latency, and NPU/GPU budgets are real.
Why local AI still matters
Local inference is not only a privacy story. It is a systems story:
- Latency and offline resilience
- Cost control under fixed hardware
- Product features that must work next to the user
- Alignment with hardware-optimized stacks (Lemonade, Ryzen AI, edge deployment)
If you care about on-device LLMs, local AI agents, or open-source interview tooling, the Interviewer repo and Lemonade Server path are concrete starting points.
Related on this site
- Projects — Lemonade, NPU kits, research code
- Publications — Universal Conditional Logic
- About — experience timeline (including ex-AMD)
- Contact — how to reach me
I validate this class of work with unit, integration, and performance checks on production-like client hardware—because local AI that only works in a lab is not local AI users can trust.