Speaker
Description
VoxBento: Building an Open Source Real-Time Multilingual Interpretation Platform
Language barriers continue to limit participation in conferences, classrooms, and global communities. While commercial interpretation platforms exist, they are often expensive, proprietary, and difficult to self-host.
VoxBento is an open source real-time multilingual interpretation platform developed under FOSSASIA during Google Summer of Code 2026 to make AI-powered live interpretation accessible to everyone.
This talk takes the audience through the engineering behind VoxBento, from capturing live audio with WebRTC to speech recognition, machine translation, text-to-speech synthesis, and streaming interpreted audio back to listeners with minimal latency.
Topics covered include:
- Designing a low-latency streaming pipeline
- WebRTC (WHIP/WHEP) architecture for real-time communication
- Integrating multiple AI providers for speech recognition, translation, and text-to-speech
- Building provider-agnostic and self-hostable infrastructure
- Lessons learned while developing an open source project during GSoC
- Future directions for open, AI-powered accessibility tools
The session concludes with a live demonstration of VoxBento and practical ways developers can contribute to the project.
Audience: Developers, students, AI enthusiasts, and open source contributors interested in distributed systems, real-time communication, and applied machine learning.
No prior knowledge of speech AI is required.
Session author's bio
I'm Arnav Angarkar, a Computer Science student at IIIT Dharwad and an open source contributor with FOSSASIA. My work primarily focuses on AI infrastructure, developer tools, and scalable web applications.
As a Google Summer of Code 2026 contributor, I developed major components of VoxBento, an open source real-time multilingual interpretation platform that combines speech recognition, machine translation, and speech synthesis for live events. I also actively contribute to Eventyay, FOSSASIA's open source event management platform, where I've worked across both backend and frontend systems.
I'm passionate about building practical open source software that improves accessibility and enables communities worldwide to collaborate across language barriers. I enjoy sharing engineering lessons from real-world projects and helping new contributors get involved in open source.
Any other info we should know?
This is a technical talk aimed at developers interested in AI, distributed systems, and open source.
Happy to provide a live demonstration during the session. An offline recording will be available as backup.
The audience will learn:
How real-time multilingual interpretation systems work
Designing low-latency streaming architectures using WebRTC
Integrating speech recognition, translation, and text-to-speech providers
Lessons learned while building VoxBento as a GSoC 2026 project under FOSSASIA
Opportunities to contribute to the project
The presentation includes architecture diagrams, live demonstrations, and discussion of real engineering challenges encountered during development. Basic familiarity with web development is helpful but not required.
| Please confirm that there are included headshots of all speakers in their profiles | Yes |
|---|---|
| In Person Attendance | Remote |
| Agree to Privacy Policy and Notice | I agree |
| Level of Difficulty | Intermediate |
| Social Media | https://x.com/ArnavA_0824 |