AI Agent Architect
Emergentlabsinc · Bengaluru · 2026-06-29
About this role
<p class="font-claude-response-body break-words whitespace-normal">Emergent builds autonomous coding agents that replace traditional software development by generating, testing, and deploying production applications directly from plain-language intent. Our systems run in production at global scale and are used to build millions of real applications.</p> <p class="font-claude-response-body break-words whitespace-normal">Since our public launch, we've crossed <strong>$100M in ARR</strong> and grown to <strong>over 10M users across 190+ countries</strong>. We're backed by <strong>Khosla Ventures, SoftBank, Google, Lightspeed, Prosus, Together, and Y Combinator.</strong></p> <p class="font-claude-response-body break-words whitespace-normal">We're solving the hard part of AI-driven software creation: correctness, reliability, security, and scale in real production systems. The team is built by <strong>repeat founders, Olympiad medalists, IIT &amp; IIM alumni,</strong> and leaders from <strong>Google, Amazon, and Dropbox.</strong></p> <p class="font-claude-response-body break-words whitespace-normal">We're hiring builders who want ownership, speed, and impact at global scale.</p> <p class="font-claude-response-body break-words whitespace-normal"><strong>The Role:</strong><br>We're building AI agents that plan, build, test, and ship real software at global scale. Your job is to turn raw LLM and system capability into measurable, shippable gains in how well our agents actually perform. You own the loop end-to-end: what we try, how we prove it's better, what goes live, and what gets rolled back. This role sits at the intersection of product, engineering, and applied research. You won't just ship features. You'll decide what makes the agent better, how we measure it, and what's ready to go live at the scale of millions of real applications.</p> <p class="font-claude-response-body break-words whitespace-normal"><strong>What You'll Do:</strong></p> <ul class="[li_&amp;]:mb-0 [li_&amp;]:mt-1 [li_&amp;]:gap-1 [&amp;:not(:last-child)_ul]:pb-1 [&amp;:not(:last-child)_ol]:pb-1 list-disc flex flex-col gap-1 pl-8 mb-3"> <li class="font-claude-response-body whitespace-normal break-words pl-2">Develop deep, evidence-grounded intuition for how the agent thinks, succeeds, and fails across the full range of real-world use</li> <li class="font-claude-response-body whitespace-normal break-words pl-2">Mine production behavior for the failure modes, regressions, and bottlenecks most teams never see, and turn them into clear, quantitative signal</li> <li class="font-claude-response-body whitespace-normal break-words pl-2">Define and run high-leverage experiments that improve agent quality, reliability, and code outcomes, spanning prompt tuning, eval dataset creation, experimentation, and harness engineering</li> <li class="font-claude-response-body whitespace-normal break-words pl-2">Build and evolve evaluation frameworks that measure agent quality at scale: define the metric, build the dataset, validate it against known signals, and ship dashboards that make regressions impossible to miss</li> <li class="font-claude-response-body whitespace-normal break-words pl-2">Ship with rigor through clear metrics, evaluation gates, staged rollouts, and explicit rollback criteria</li> <li class="font-claude-response-body whitespace-normal break-words pl-2">Drive the hard problems in context engineering, memory systems, tool use, and long-horizon execution, where agent reliability is actually won or lost</li> <li class="font-claude-response-body whitespace-normal break-words pl-2">Make hard calls in subjective, probabilistic systems: when a regression is real, when a win is noise, when a benchmark is overfit, when to ship on mixed signals, and when to kill a promising direction</li> <li class="font-claude-response-body whitespace-normal break-words pl-2">Think like the agent and continuously make it smarter, more reliable, and more useful</li> </ul> <p class="font-claude-response-body break-words whitespace-normal"><strong>Who You Are:</strong></p> <ul class="[li_&amp;]:mb-0 [li_&amp;]:mt-1 [li_&amp;]:gap-1 [&amp;:not(:last-child)_ul]:pb-1 [&amp;:not(:last-child)_ol]:pb-1 list-disc flex flex-col gap-1 pl-8 mb-3"> <li class="font-claude-response-body whitespace-normal break-words pl-2">5+ years building and shipping software, with real end-to-end ownership</li> <li class="font-claude-response-body whitespace-normal break-words pl-2">At least 1 year of hands-on experience with agentic systems and their evaluation</li> <li class="font-claude-response-body whitespace-normal break-words pl-2">Strong systems thinking paired with sharp product judgment</li> <li class="font-claude-response-body whitespace-normal break-words pl-2">Hands-on with experimentation, metrics, and debugging, fluent in Python, SQL, or similar for analysis</li> <li class="font-claude-response-body whitespace-normal break-words pl-2">Comfortable reasoning about noise, confounds, distribution shift, and selection effects, you know when to trust a number and when to suspect it</li> <li class="font-claude-response-body whitespace-normal break-words pl-2">Energized by the long tail, sifting through large volumes of agent behavior to find the rare, hidden failure mode is the fun part for…
Skills asked for
- llm
- go
- python
Your next role is already in here.
Search live openings from thousands of employers, save the ones worth a second look, and let JobBob keep watch for the rest.