axiombio - Agent Harness Engineer
Responsibilities
• Own the harness: the scaffolding, tooling, and infrastructure that turn frontier models into agents that do long-horizon scientific analysis • Build the data backbone that gets everything to the right place: pipelines, storage, and systems for runtime context, agent trajectories, eval results, and training data • Build sandboxed execution environments with instant spin-up/tear-down, reproducible and deterministic enough to trust for evals and RL • Design and run the eval systems: offline suites, test cases on production traces, LLM-as-judge pipelines, regression gates • Work closely with domain experts and encode their taste into rubrics, golden sets, and review workflows, turning "I know it when I see it" into something measurable • Build the tools the agent needs to do better work, and stay tuned in to the state of the art for new methods, protocols, and patterns worth adopting • Make every agent run observable and replayable: trace every model call, tool call, and state transition, and build the debugging tooling to make sense of it • Engineer the context: memory, compaction, retrieval, and recovery so long-horizon agent runs stay coherent across hours and crashes • Own the loop itself: retries, budget caps, stop conditions, output verification, permissions, and guardrails • Support ML research with environments, reward instrumentation, and rollout infra for RL on agentic tasks • Python, Modal, DuckDB, FastAPI, Docker, Containerization, Terraform • Engineers who've built with LLM APIs and shipped agentic systems: tool use, loops, and the debugging scars to prove it • Built bespoke evaluation, monitoring, and RL env observability tooling (SvelteKit, Svelte 5, React) • What we look for: • Can tackle deep technical challenges and own/ship simple, clean, maintainable code • High ownership: owns outcomes end to end, not tickets, and doesn't wait for a spec to start moving • Allergic to complexity: reaches for the simplest system that works and keeps it that way as it scales • Strong software engineer first, with infrastructure, platform, data, or devtools depth and production systems they're proud of • Instinctively asks "how would we know if this is working?" and builds the measurement alongside the feature • Reads an agent failure trace the way other engineers read a stack trace • Cares about reliability because they know environments that break silently poison evals and training data • Has a knack for surfacing the important questions about what the agent actually needs to do better work • Sharp and confident • keeps up with a really technical & sophisticated crowd • comfortable getting in over their head and figuring it out as they go • thrives in a discipline with no playbook, because it's being invented right now • Curious about how things work: an engineering/tinkering mindset, good at scavenging the state of the art • Passion for learning what "good" looks like from deep domain experts and turning it into systems
Apply in one click
Upload My Resume
Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT