Automated Web-Based Voice Counseling Systems: A Client-Side Architecture For Low-Latency Institutional Dialogue
Keywords:
Web Speech API, Large Language Models, Voice Automation, Human-Computer Interaction, Asynchronous Concurrency, Client-Side Speech Processing.Abstract
Purpose: Institutions that handle large volumes of repetitive inbound enquiries, such as universities, hospitals and public service departments, face a persistent operational dilemma. Staffing a call centre with trained human counsellors is expensive and difficult to scale during peak admission or enrolment cycles, while button-driven Interactive Voice Response (IVR) systems are cheap to run but routinely frustrate callers who simply want a direct, conversational answer. This paper sets out to resolve that dilemma by describing a voice counselling system that runs totally inside a standard web browser, requires no application installation, and can hold a natural spoken conversation with a prospective student or patient without depending on costly telephony infrastructure.
Design/Methodology: The study adopts an applied systems-design methodology. Rather than testing a single algorithm in isolation, the paper documents the complete engineering pipeline of a working prototype: the browser-native speech interfaces that capture and reproduce sound, the finite state controller that governs who is allowed to speak at any given moment, the prompt-engineering rules that keep a Large Language Model's replies short enough to sound like a real phone call, the mathematical model used to animate an on-screen waveform, and the security layer that a production deployment would require. Each component is described together with the reasoning that led to its particular design, and the resulting architecture is compared against conventional server-side voice bot deployments.
Findings: Coupling the native Web Speech API directly to a remote Large Language Model, lacking any intermediate streaming server, removes the largest source of latency found in conventional cloud telephony pipelines. The browser performs all audio capture and audio playback locally, sends only plain text across the network, and receives only plain text in return. A software mutex implemented purely in JavaScript is sufficient to prevent the system's own voice from being picked up by its own microphone, which is normally one of the hardest problems in real-time voice engineering. A small trigonometric function can additionally produce a waveform animation that looks convincingly alive without ever requesting access to the raw microphone amplitude stream, which keeps the permission footprint of the application small.
Practical Implications: The architecture described here gives institutions a way to deploy a voice counsellor that a caller can reach from any modern browser tab, with no app download, no dialed access number and no per-minute telephony bill. Section 6 of the paper also sets out the minimum security posture, namely a server-side reverse proxy and structured error recovery that any team wishing to move the prototype into production should adopt before exposing it to the public internet.
Originality/Value: Most published work on conversational voice agents still assumes a server-mediated audio pipeline built on WebRTC or a telephony gateway such as Twilio. This paper instead documents a fully client-side alternative, contributes a mutex-based technique for blocking acoustic feedback without hardware echo cancellation, and proposes a zero-permission mathematical model for visualizing speech that avoids the privacy and battery costs of reading the live microphone stream.





