Why Voice AI Needs to Speak Our Slang (And How to Build One That Doesn't Break)
Typing prompts into a chatbot when you're trying to learn something hard is painful. Here is a breakdown of low-latency voice pipelines, code-switching, and why audio tutors actually make sense.

Nobody wants to type out a three-step calculus problem on a cracked phone screen while sitting in a noisy bus park or rushing before the inverter battery gives up.
Text interfaces are fine for search queries, but for learning, they suck the life out of the interaction. You spend half your energy trying to craft a neat prompt instead of just asking the question floating around in your head.
I was recently looking at a project called Bharat Buddy, built by an engineer during a 10-day voice AI challenge. The creator set out to solve a specific headache for students in India: building a voice tutor that understands Hindi, English, and Hinglish natively, remembers context, and talks back fast enough to feel like a real conversation.
Looking at the architecture, the lessons apply word-for-word to what we need to build for our own market.
The Latency Trap in Real-Time Voice
Most developers think building a voice agent is just chaining three APIs together: Audio In → Whisper API → OpenAI ChatCompletion → ElevenLabs TTS → Audio Out.
If you build it that way, your user will say "Hello" and wait five awkward seconds in dead silence before the bot replies. That delay instantly breaks the illusion. Nobody talks like that.
To make voice AI feel human, the round-trip latency needs to sit under 800 milliseconds.
The stack in the Bharat Buddy build handled this using LiveKit for low-latency WebRTC transport, combined with Murf Falcon for fast text-to-speech generation. LiveKit manages the bidirectional audio stream without the overhead of standard HTTP polling or clunky WebSocket handshakes. The moment the user stops speaking, the speech-to-text pipeline triggers the LLM, which streams tokens directly into the TTS engine chunk by chunk.
You do not wait for the whole paragraph to finish generating before you start streaming audio back to the user's ears. That is how you kill the lag.
Code-Switching is Mandatory, Not a Feature
What caught my eye in this build was the focus on Hinglish—the seamless blend of Hindi and English.
Here in Nigeria, nobody speaks textbook Queen’s English when they are stuck or stressed. If a student in a federal university hostel in Akure or a secondary school kid in Enugu is confused about physics, they won't say: "Could you kindly elucidate the second law of thermodynamics?"
They will say: "Abeg, explain this force and momentum thing, I no get the formula."
If your model falls flat on its face whenever someone mixes Nigerian Pidgin, Yoruba, or Igbo phrases with standard English, you haven't built an accessible product. You've built a toy for people who already have MacBook Pros and private tutors.
Building an agent that handles colloquial code-switching means picking an LLM fine-tuned on regional dialects, or setting system prompts that explicitly normalize the blend without scolding the user.
Memory Without Token Bloat
The developer behind Bharat Buddy made a practical engineering choice around memory: do not try to feed the whole raw conversation history back into the LLM on every turn.
If you store every single audio transcript chunk, your context window explodes, your latency goes through the roof, and your API bill will give you hypertension.
Instead, practical memory should be structured:
- User profile state: What topics has this person struggled with previously?
- Current session goals: Are we working on fractions or quadratic equations right now?
- Session summaries: Compress the last 10 turns into a 2-sentence summary and feed only that into the system prompt.
Keep the memory lean. The agent only needs to know enough to not ask "What's your name?" every two minutes.
The Escape Hatch: Tools and Human Escalation
The biggest mistake founders make with AI agents is pretending the bot knows everything.
In this voice agent build, the LLM had clear tool-calling boundaries. Need to solve a math problem? Don't let the LLM guess the arithmetic—hand it off to a deterministic calculation tool. Is the student getting frustrated or asking something outside the safety guardrails? Trigger an outbound escalation hook to a human tutor.
An AI voice agent shouldn't be a closed loop. It should be the front desk that handles 80% of the repetitive queries instantly, and gracefully hands off the remaining 20% to specialists or external tools.
If you are hacking on voice tech right now, stop thinking about generic chatbots. Pick a specific, noisy problem, optimize your audio streaming pipeline for speed, and make sure the thing actually understands how real people speak.
Related from Engineering
Let's build your next big product.
Accepting project-based freelance, remote engineering roles, and hybrid positions.