Deepgram
Embedded AI Engineer, On-Device Models
Full-time00+ yrsUSA | RemoteRemoteNot disclosedApply by 5 Sept 2026
Overview
COMPANY OVERVIEW Deepgram is the leading platform underpinning the emerging trillion-dollar Voice AI economy, providing real-time APIs for speech-to-text (STT), text-to-speech (TTS), and building production-grade voice agents at scale. More than 200,000 developers and 1,300+ organizations build voice offerings that are ‘Powered by Deepgram’, including Twilio, Cloudflare, Sierra, Decagon, Vapi, Daily, Cresta, Granola,
What you'll do
- They'll run directly on the small, low-power devices people carry, wear, and keep around their homes: phones, earbuds, wearables, appliances, cameras, and purpose-built consumer hardware.
- Putting state-of-the-art speech models on devices with tight memory, compute, thermal, and battery budgets is a fundamentally different engineering problem, and it's one of the most important frontiers for bringing voice AI to everyone.
- As an Embedded AI Engineer, you will take Deepgram's models and make them run, fast, accurately, and efficiently, on resource-constrained embedded and edge platforms.
- You'll work across the stack: optimizing and compiling models for on-device inference, writing performance-critical runtime code, and squeezing every last millisecond and milliwatt out of a wide range of mobile application processors, embedded SoCs, microcontrollers, and dedicated AI accelerators.
- Your work directly enables a new class of private, offline-capable, real-time voice experiences on the devices closest to the user.
- This role is a great fit whether you're a hands-on senior embedded engineer who wants to go deep on a hard problem, or a staff-level technical leader who wants to define how Deepgram's voice AI gets onto consumer hardware and raise the bar for the engineers around you.
- We'll set the level to your experience.
- Take Deepgram's Speech and Conversational models and get them running on embedded and low-power consumer hardware, defining the architecture for on-device, real-time inference across a diverse range of processors and accelerators.
Requirements
- Every team member who works at Deepgram is expected to actively use and experiment with advanced AI tools, and even build your own into your everyday work.
- We measure how effectively AI is applied to deliver results, and consistent, creative use of the latest AI capabilities is key to success here.
- Candidates should be comfortable adopting new models and modes quickly, integrating AI into their workflows, and continuously pushing the boundaries of what these technologies can do.
- Additionally, we move at the pace of AI.
- Change is rapid, and you can expect your day-to-day work to evolve just as quickly.
- This may not be the right role if you’re not excited to experiment, adapt, think on your feet, and learn constantly, or if you’re seeking something highly prescriptive with a traditional 9-to-5.
Skills
GoRustRAGC++