Backend Developer — Voice AI & SIP/VoIP
FabTechSol · Pakistan
Job Description
Backend Developer Voice AI – Real-Time SIP/VoIP + Urdu-English Speech Pipeline
The Core Technical Challenge
Why this is hard and why it matters
International VoIP providers – Twilio, Vonage, Plivo – have no support for Pakistani phone numbers. That means any production voice AI system for this market must be built from the ground up using local SIP trunks and GSM gateways. There is no off-the-shelf solution.
On top of the infrastructure problem, Pakistani urban speech is heavily code-switched – callers naturally mix Urdu and English mid-sentence (Urdish). Standard STT models trained on monolingual audio perform poorly. The pipeline must handle this gracefully in real time, with a sub-2.5s end-to-end latency from the caller speaking to hearing a response.
What You'll Build
- Whisper STT processing Urdu-English mixed audio reliably under 1 second
- GPT-4o conversation engine with structured output via function calling
- Graceful handling of silence, interruptions, and unclear input mid-call
- Per-call context loaded dynamically from database — changes reflect within 5 minutes
- Conversation data saved to PostgreSQL and pushed to dashboard via WebSocket in real time
- Call recording + full transcript stored and accessible per session
- WhatsApp Business API notification on key events
- Multi-tenant management dashboard with real-time updates, call logs, and session replay
Core Backend
- Python FastAPI – async, production-grade
- OpenAI Whisper (whisper-1) – Urdu/English mixed-language STT
- OpenAI GPT-4o – LLM with function calling / structured output
- Google Cloud TTS (ur-PK) or ElevenLabs – natural Urdu-English TTS
- Sub-2.5s end-to-end latency from speech to response
SIP / Call Infrastructure
- FreeSWITCH – SIP server, inbound call routing, audio streaming
- GSM Gateway (GoIP) or PTCL SIP Trunk integration
- Call recording, transcript storage, real-time audio relay
- Concurrent call handling (5+ simultaneous, 50+ at scale)
- React-based management dashboard with real-time WebSocket updates
- WhatsApp Business API notifications (Twilio/360dialog)
- Clean, presentable UI for client demos and onboarding
Infrastructure & Ops
- Ubuntu VPS deployment (local PK or DigitalOcean)
- Low-latency audio hosting – local Pakistani VPS preferred
- Multi-tenant architecture with strict data isolation
- 99.5%+ uptime reliability design
- Urdu-English code-switching (Urdish) conversation handling
- Pakistan Standard Time (PKT, UTC+5) throughout
- Familiarity with Pakistani telecom (Jazz, Telenor, PTCL, Ufone, Zong)
- Natural-sounding Pakistani voice persona for TTS output
This is a key deliverable, not optional
- Streams audio to the FastAPI backend for Whisper processing
- Displays conversation transcript in real time
- Loads current configuration and context for the demo session
- Clean, presentable UI – this is shown to clients
Apply for This Role
Application Form
Fill out the form below. We review every application carefully. Only candidates with verifiable voice/VoIP/audio experience will be considered.
Personal Information
Experience & Engagement
Years of Experience *
Preferred Engagement *
Work Arrangement *
Availability *
Technical Details
Key Skills * (comma-separated)
LinkedIn Profile
Voice AI & VoIP Specific Experience *
Upload CV (PDF or Word, max 10MB)
Experience Confirmations *
All six must be confirmed to be considered. Check every item that applies to your actual experience.
- I have built real-time audio streaming applications in production
- I have integrated OpenAI Whisper (STT) or similar speech-to-text in a live system
- I have worked with SIP/VoIP infrastructure (FreeSWITCH, Asterisk, or equivalent)
- I have built WebSocket-based systems handling concurrent real-time connections
- I have built and deployed Python FastAPI applications at production scale
Details
| Company | FabTechSol |
| Location | Pakistan |
| Type | FULL TIME |
| Niche | tech |
