Agent Infrastructure Engineer — Core Harness (Superagent)
Imagineart · India · 2026-08-19
About this role
About ImagineArt
We're redefining how the world creates and designs.
ImagineArt is one of the fastest-growing GenAI companies in the world. We've scaled faster than most funded startups — with zero outside funding.
• $35M+ ARR crossed this year
• 100M+ social impressions
• Built and shipped our own image generation model, now ranked #3 globally for photo realism
No funding. No shortcuts. Just a sharp, driven team building one of the strongest GenAI products in the world — and we're just getting started.
We're looking for an Agent Infrastructure Engineer to own Superagent, our core agent harness that powers conversations, tool calls, and multi-step agentic workflows across our AI products.
This is a deep systems and infrastructure role — not prompt engineering and not simply wrapping model APIs. You'll work on the core orchestration loop, tool-calling infrastructure, context and memory management, streaming, retries, evaluation, observability, and performance.
Key Responsibilities
• Own the architecture, development, and evolution of Superagent, our core agent harness.
• Design and optimize the agent execution loop for latency, reliability, token efficiency, cost, and task completion.
• Build and improve core harness systems including context management, memory/state handling, tool routing, function schemas, structured outputs, retries, and error recovery.
• Build and maintain agent evaluation infrastructure to measure quality and guide engineering decisions with data.
• Integrate and benchmark multiple LLM providers and models, evaluating performance, cost, reliability, and capabilities.
• Implement performance optimizations such as caching, batching, parallel tool execution, and prompt/context compression.
• Build deep observability and instrumentation across agent runs, including tracing, logging, metrics, and regression detection.
• Extend and customize underlying agent frameworks when existing abstractions are insufficient.
• Build reliable integrations with evolving AI and tool ecosystems.
• Work closely with product engineering teams to expose clean abstractions while keeping harness complexity behind the platform.
• Debug and resolve complex issues across non-deterministic, distributed, and model-driven systems.
Required Skills & Qualifications
• 4+ years of experience in software engineering, backend engineering, or systems infrastructure.
• Strong proficiency in Python and/or TypeScript.
• Hands-on experience building or operating LLM-based agents in production.
• Strong understanding of tool calling, function schemas, context limits, structured outputs, model failures, and unreliable LLM behavior.
• Experience with at least one agent framework such as LangGraph, OpenAI Agents SDK, CrewAI, AutoGen, or a custom/homegrown agent harness.
• Strong understanding of agent orchestration and multi-step workflows.
• Experience building or working with evaluation suites, benchmarks, A/B testing, or other measurement systems for AI products.
• Strong understanding of concurrency, caching, profiling, performance optimization, and latency/cost tradeoffs.
• Experience working with LLM APIs and production AI infrastructure.
• Excellent debugging and problem-solving skills, especially for complex and non-deterministic systems.
• Passionate about technology, self-driven, and proactive with a strong builder mindset.
Optional / Nice-to-Have Skills
• Contributions to open-source agent frameworks, LLM tooling, or AI infrastructure.
• Experience with RAG pipelines, vector databases, or long-term memory systems for AI agents.
• Familiarity with MCP (Model Context Protocol) or similar tool-integration standards.
• Experience with LLM inference infrastructure, model routing, rate limits, fallbacks, or high-volume model APIs.
• Experience with LangChain, LlamaIndex, LangGraph, DSPy, or similar AI infrastructure frameworks.
• Experience with Kubernetes, Docker, cloud infrastructure, or distributed systems.
• Experience building internal developer platforms or infrastructure used by multiple engineering/product teams.
• Strong background in observability, distributed tracing, and production reliability.
• Contributions to open-source projects or personal AI infrastructure projects.
Why Join Us?
• Own the core agent infrastructure behind our AI products — every improvement you make can multiply across the entire platform.
• Work on real production-scale AI systems, not demo agents or simple API wrappers.
• Solve challenging problems across LLMs, distributed systems, orchestration, performance, and infrastructure.
• Have direct influence over the architecture and technical roadmap of our entire agent stack.
• Collaborate with a passionate and talented team building some of the most ambitious GenAI products in the market.
• Competitive salary and benefits package.
• A culture that encourages ownership, experimentation, learning, and data-driven engineering.
Originally posted on Himalayas
Skills asked for
- llm
- python
- typescript
- kubernetes
- docker
Your next role is already in here.
Search live openings from thousands of employers, save the ones worth a second look, and let JobBob keep watch for the rest.