Senior Multi-Agent Systems Engineer
Role Overview
We are building a new generation of multi-agent AI systems for complex knowledge work and long-horizon autonomous tasks. Having completed early prototype validation, we are now entering the platformization and production engineering phase.
We are seeking a senior engineer to own core multi-agent runtime and platform infrastructure. You will lead architectural refactoring of our existing prototype—designing unified agent abstractions, task scheduling, state and memory layers, tool governance, and observability systems—to enable complex multi-agent workflows to run reliably, evolve continuously, and be debugged, reproduced, and extended at scale.
This is not a prompt engineering role, nor one that simply wires together existing agent frameworks. We value systems engineering judgment, production experience with multi-agent collaboration, and a keen sensitivity to failure modes in long-running distributed systems.
Key Responsibilities
- Architectural Refactoring: Lead architectural refactoring of the multi-agent platform, crystallizing prototype-era implicit couplings and ad-hoc conventions into a unified, extensible platform architecture with clear boundaries.
- Agent Abstractions & Runtime Models: Design unified agent abstractions and runtime models, including lifecycles, message handling, state transitions, tool invocation, model provider adaptation, budget control, failure recovery, and exit semantics.
- Scheduling & Collaboration: Build multi-agent scheduling and collaboration mechanisms supporting parallel execution, task delegation, result aggregation, conflict resolution, and long-horizon workflow evolution.
- State, Memory & Tool Governance: Design cross-agent structured state and memory layers with support for versioning, concurrency safety, atomic commits, schema evolution, and crash recovery; establish tool governance covering schemas, sandboxing, audit, error handling, and external protocol integration.
- Observability & Reproducibility: Build observability and reproducibility for complex runs—structured logs, tracing, artifact tracking, run replay, cost attribution, and failure taxonomy—pushing systemic failure modes to be resolved at the architectural layer rather than patched at the application layer.
Required Qualifications
- 5+ years of software engineering experience with strong backend, platform, or distributed systems background; 2+ years building LLM agents, multi-agent systems, workflow engines, or complex asynchronous task systems.
- Led or deeply participated in architectural refactoring or production reliability transformations of medium-to-large systems, with ability to articulate tradeoffs and outcomes.
- Deep Python expertise with strong grasp of asyncio, concurrency models, cancellation semantics, timeout control, resource cleanup, error propagation, and system-level debugging.
- Familiarity with at least one major agent/workflow/orchestration framework (LangGraph, AutoGen, Temporal, Ray, Prefect, Airflow, etc.) and hands-on understanding of real LLM/Agent failure modes: format errors, tool failures, context drift, retry loops, cost runaway, and state contamination.
- Experience designing tool invocation systems (function calling, schemas, permissions, sandboxing, audit) and state management (transactions, idempotency, atomic writes, crash recovery, schema migrations); track record building observability systems (tracing, structured logs, metrics, run replay).
- Strong engineering aesthetic: bias toward architectural solutions over quick patches; ability to write clear design docs that decompose ambiguous problems into verifiable engineering plans.
Preferred Qualifications
- Production experience building multi-agent systems, AI coding agents, agent runtimes, or LLM workflow engines; deep source-level familiarity with mainstream agent framework internals and design tradeoffs.
- Experience with MCP or similar tool/context protocols; background in complex task orchestration, distributed workers, checkpoint/resume, event sourcing, or replay debuggers.
- Experience with OpenTelemetry or agent tracing platforms (LangSmith, etc.); familiarity with Docker, Kubernetes, Ray, GPU scheduling, container isolation, or untrusted code sandboxing.
- Security awareness around prompt injection, tool injection, data exfiltration, and secret management; formal thinking around system invariants, state machines, and failure semantics; open-source contributions or public technical writing in agent/LLM infrastructure.

