Base Career helps you apply smarter for this job.
At Bland.com , our mission is to empower enterprises to build AI phone agents at scale. Voice is quickly becoming the primary interface between businesses and their customers, and we are building the models and infrastructure that make those interactions feel natural, reliable, and genuinely human.
We’ve raised $65M from leading investors including Emergence Capital, Scale Venture Partners, Y Combinator, and founders of Twilio, Affirm, and ElevenLabs.
We are looking for someone to spearhead the development of our next-generation multimodal LLM stack, combining speech, text, tools, and real-time reasoning into a single unified system. You’ll be responsible for building industry-leading conversational AI models that power Bland's agent, and taking them all the way from idea to production.
At Bland, we're not just thinking about text modeling. You will define how our agents listen, think, and act in real time , integrating streaming audio, tool execution, and dynamic context into a single coherent system.
This role sits at the intersection of:
LLM architecture and fine-tuning
real-time speech systems
agent design (prompting + tools + policies)
multimodal reasoning (audio + text + actions)
You will take ideas from research through production systems serving millions of calls per day.
Experience with LLMs, multimodal models, or speech-language systems
Deep understanding of prompting, fine-tuning, and alignment techniques
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
San Francisco, USA
San Francisco, USA
San Francisco, USA
San Francisco, USA
San Francisco, USA
San Francisco, USA
Familiarity with streaming or real-time inference is a strong plus
Ability to reason about full systems, not just models
Comfortable designing interactions between: model tools prompts runtime constraints
model
tools
prompts
runtime constraints
You can go from idea → dataset → experiment → conclusion in days
You know how to design experiments that actually answer the question
Strong sense for what makes an interaction feel natural vs robotic
Ability to translate abstract modeling ideas into user-facing improvements
You take ownership from research through deployment
You thrive in ambiguous, fast-moving environments
You care about impact, not just elegance
You think in systems, not just models
You obsess over latency, correctness, and real-world behavior
You are comfortable discarding ideas quickly when data disagrees
You push toward simple abstractions for complex problems
Experience with real-time voice systems or conversational AI
Background in tool-using agents or agent frameworks
Experience with multimodal datasets (audio + text + actions)
Contributions to LLM or speech-related research or open source
Your work will define how our agents:
understand users in real time
decide when to respond
choose what tools to call
balance speed vs correctness
behave under complex policies
This is the core intelligence layer of the product.
Verified company details for this employer are not available yet.
USD 140000-250000 yearly / year
Full-time
Mid
Remote
Apply faster on company sites with our extension.