Dinesh/ Blog
← All articles
cloud engineering

AI Voice chatting to help with Customer support use-cases

Using current speech augmented LLMs (SpeechLLMs) with realtime voice modality to understand user issues and to provide support and solutions.

5 min readChecking device speech…
Lesson preparation & details

Level: beginner

Application engineering projects · Lesson 4 of 6 ↗
Loading this browser’s progress…

By the end, you should be able to

  • Transcription succeeds but synthesis times out. Should the whole job start again and recreate the ticket?
  • Explain the version and execution boundaries before applying the examples

Bring with you

  • Basic programming and HTTP; follow the chapter or cloud-track sequence

Editorial review: · What review means

In this article · 15 sections

Review and execution boundary

Browser MediaRecorder and staged STT→LLM→TTS architecture; recording demo is not a reproduced realtime conversation system.

Reviewed on 7 October 2026 against the official source snapshots linked below. The review is bounded editorial correction, not certification of every dependency, security property or cloud deployment. Historical setup commands and optional exercises were not executed. No cloud resources, third-party packages or external side effects were created. Old screenshots and unavailable private assets remain in the private recovery archive, not prerequisites for this lesson.

Using Realtime speech LLMs to assist with service requests

Provide solutions and raise service tickets.

The goal is a support conversation that starts with a spoken request and eventually returns an answer or raises a ticket. This first installment stops earlier: it records and deploys the audio-capture web app. Transcription, model responses, speech playback and ticket creation belong to the planned system, not the completed demo.

Initial Architecture - Version V1

The proposed architecture is a staged speech-to-text, text-model and text-to-speech pipeline rather than a demonstrated realtime speech model. Follow one request through the diagram: recorded audio becomes text, the language model generates a reply, and the reply is converted back to audio. Each handoff adds work and latency that the recorder alone cannot measure.

Link to the architecture diagram : User draw.io to render

  • Using webapp get the audio input from user with record button from UI
  • Pass to a audio to text conversion tool (google or other services)
  • Use LLM model - Open AI/Gemini to get the response back for the converted text
  • Use text to speech to convert back the text generated from the Large language model
  • Return the audio file to webapp

This is the target flow for version 1. Conversation history and service-ticket creation would add state and external actions beyond that flow; neither is implemented in the recording work documented below.

Part 1 : Recording audio from user and generating a audio file

Using replit agent created a voice recording flask app which we can leverage

  • Setting up the code base locally to test and to deploy this as webapp

  • Pushing the code to the repo : link

  • Creating a webapp and resource group to deploy and run this app in azure : techvistara-ai-voice.azurewebsites.net (deployment endpoint unavailable as of 2026-10-07)

  • Enabling the deployment and attaching to the above repo, workflow yml : yml

  • Following the initial blog steps to add startup command in app portal and configuring the app service to run the flask application

gunicorn --bind 0.0.0.0:$PORT main:app
  • Successfully deploying the app to app service using github actions

  • Verified the deployment by accessing the app service link, the web server is running successfully

Historical deployment endpoint

https://techvistara-ai-voice.azurewebsites.net/ was the endpoint used for the deployment shown below. It is unavailable as of 7 October 2026; the screenshots document the original deployment, not a currently running service.

Flow diagram until completed portion

Compare this progress diagram with the target architecture above. The completed portion establishes audio capture and web hosting; it does not yet show that a spoken request can be transcribed, answered or turned into a ticket. That boundary explains why speech-to-text is the next step.

Next steps

Generating text from the generated audio file from the user

  • Exploring Azure Speech to Text service

Project Suspension Notice

Due to budget constraints and the need to prioritize other projects, I have decided to temporarily suspend the AI Voice chatting application. The web application will be shut down until the next steps are designed and finalized. All progress has been archived and can be accessed here: Project Archive

Corrected contracts and failure analysis

Recorded-audio upload is a batch pipeline, not automatically full-duplex realtime speech. Negotiate an actually supported recording MIME/codec, cap duration/bytes and release microphone tracks when the user stops or navigates away. Ask for consent and provide a visible recording indicator; transcripts and recordings need an explicit retention/access policy.

Model each stage as a job with a stable ID: captured, uploaded, transcribed, answered, synthesized or failed. Preserve language and confidence metadata without claiming a confidence score proves correctness. Ticket creation is an external side effect: require user confirmation, validate extracted fields and use an idempotency key. A repeated callback must not create two tickets. Budget latency separately for upload, recognition, retrieval/model, synthesis and playback; the recorder alone measures none of the downstream stages.

Boundary exercise with solution

Transcription succeeds but synthesis times out. Should the whole job start again and recreate the ticket?

Solution and reasoning

No. Retain the completed text response and confirmed ticket ID, retry only the failed stage with bounded attempts, and provide text fallback. Track stage results under the same job ID and avoid storing raw audio beyond the retention purpose.

Source-backed review notes

Pause / Recall / Apply

Can you explain it without the page?

Close the example. Reconstruct the core idea, then change one assumption. Mark complete when you’re ready; you can always undo it.

Stored in this browser only. No account, no sync. Clearing browser data removes your record.