AI Voice chatting to help with Customer support use-cases
Using current speech augmented LLMs (SpeechLLMs) with realtime voice modality to understand user issues and to provide support and solutions.
Lesson preparation & details
Level: beginner
By the end, you should be able to
- Transcription succeeds but synthesis times out. Should the whole job start again and recreate the ticket?
- Explain the version and execution boundaries before applying the examples
Bring with you
- Basic programming and HTTP; follow the chapter or cloud-track sequence
Editorial review: · What review means
In this article · 15 sections
Review and execution boundary
Browser MediaRecorder and staged STT→LLM→TTS architecture; recording demo is not a reproduced realtime conversation system.
Reviewed on 7 October 2026 against the official source snapshots linked below. The review is bounded editorial correction, not certification of every dependency, security property or cloud deployment. Historical setup commands and optional exercises were not executed. No cloud resources, third-party packages or external side effects were created. Old screenshots and unavailable private assets remain in the private recovery archive, not prerequisites for this lesson.
Using Realtime speech LLMs to assist with service requests
Provide solutions and raise service tickets.
The goal is a support conversation that starts with a spoken request and eventually returns an answer or raises a ticket. This first installment stops earlier: it records and deploys the audio-capture web app. Transcription, model responses, speech playback and ticket creation belong to the planned system, not the completed demo.
Initial Architecture - Version V1
The proposed architecture is a staged speech-to-text, text-model and text-to-speech pipeline rather than a demonstrated realtime speech model. Follow one request through the diagram: recorded audio becomes text, the language model generates a reply, and the reply is converted back to audio. Each handoff adds work and latency that the recorder alone cannot measure.
Link to the architecture diagram : User draw.io to render
- Using webapp get the audio input from user with record button from UI
- Pass to a audio to text conversion tool (google or other services)
- Use LLM model - Open AI/Gemini to get the response back for the converted text
- Use text to speech to convert back the text generated from the Large language model
- Return the audio file to webapp
This is the target flow for version 1. Conversation history and service-ticket creation would add state and external actions beyond that flow; neither is implemented in the recording work documented below.
Part 1 : Recording audio from user and generating a audio file
Using replit agent created a voice recording flask app which we can leverage
-
Setting up the code base locally to test and to deploy this as webapp
-
Pushing the code to the repo : link
-
Creating a webapp and resource group to deploy and run this app in azure :
techvistara-ai-voice.azurewebsites.net(deployment endpoint unavailable as of 2026-10-07) -
Enabling the deployment and attaching to the above repo, workflow yml : yml
-
Following the initial blog steps to add startup command in app portal and configuring the app service to run the flask application
gunicorn --bind 0.0.0.0:$PORT main:app
-
Successfully deploying the app to app service using github actions
-
Verified the deployment by accessing the app service link, the web server is running successfully
Historical deployment endpoint
https://techvistara-ai-voice.azurewebsites.net/ was the endpoint used for the deployment shown below. It is unavailable as of 7 October 2026; the screenshots document the original deployment, not a currently running service.
Flow diagram until completed portion
Compare this progress diagram with the target architecture above. The completed portion establishes audio capture and web hosting; it does not yet show that a spoken request can be transcribed, answered or turned into a ticket. That boundary explains why speech-to-text is the next step.
Next steps
Generating text from the generated audio file from the user
- Exploring Azure Speech to Text service
Project Suspension Notice
Due to budget constraints and the need to prioritize other projects, I have decided to temporarily suspend the AI Voice chatting application. The web application will be shut down until the next steps are designed and finalized. All progress has been archived and can be accessed here: Project Archive
Corrected contracts and failure analysis
Recorded-audio upload is a batch pipeline, not automatically full-duplex realtime speech. Negotiate an actually supported recording MIME/codec, cap duration/bytes and release microphone tracks when the user stops or navigates away. Ask for consent and provide a visible recording indicator; transcripts and recordings need an explicit retention/access policy.
Model each stage as a job with a stable ID: captured, uploaded, transcribed, answered, synthesized or failed. Preserve language and confidence metadata without claiming a confidence score proves correctness. Ticket creation is an external side effect: require user confirmation, validate extracted fields and use an idempotency key. A repeated callback must not create two tickets. Budget latency separately for upload, recognition, retrieval/model, synthesis and playback; the recorder alone measures none of the downstream stages.
Boundary exercise with solution
Transcription succeeds but synthesis times out. Should the whole job start again and recreate the ticket?
Solution and reasoning
No. Retain the completed text response and confirmed ticket ID, retry only the failed stage with bounded attempts, and provide text fallback. Track stage results under the same job ID and avoid storing raw audio beyond the retention purpose.
Source-backed review notes
- MediaRecorder - Web APIs | MDN — accessed 2026-10-07. Exact supporting passage: “Returns the MIME type that was selected as the recording container for the MediaRecorder object when it was created.”
- How to synthesize speech from text - Speech service - Foundry Tools | Microsoft Learn — accessed 2026-10-07. Exact supporting passage: “if speech_synthesis_result.reason == speechsdk.ResultReason.SynthesizingAudioCompleted:”
- Request Files - FastAPI — accessed 2026-10-07. Exact supporting passage: “It exposes an actual Python SpooledTemporaryFile object that you can pass directly to other libraries that expect a file-like object.”
Pause / Recall / Apply
Can you explain it without the page?
Close the example. Reconstruct the core idea, then change one assumption. Mark complete when you’re ready; you can always undo it.
Stored in this browser only. No account, no sync. Clearing browser data removes your record.