Save this role or tell us whether you want more jobs like it.
We're building a system that can make and take decisions inside revenue-critical digital funnels. Not in a sandbox. In production systems where mistakes have real consequences.
The challenge is not generating answers. It is deciding what is true, what matters, and what to do about it under uncertainty. The data is messy, delayed, and often wrong. The same signal can have multiple competing explanations. There is no clean ground truth to train on. Actions can improve or break real business outcomes. Trust has to be earned before the system is allowed to act.
Most AI systems stop at insights because this part is hard. We are building systems that cross that boundary.
Today, our system already runs continuous detection and root cause analysis in production inside mission-critical enterprise workflows. The next step is building systems that can safely take action and improve these funnels over time. If we get this right, software does not just assist decision-making. It becomes part of the decision loop itself.
This is a founding engineer role for someone who wants to build real systems from first principles and ship them into production. You will work directly with the founders to design, build, evaluate, and operate the core system. That includes everything from agent behavior and retrieval patterns to backend services, product surfaces, observability, and production reliability.
This is not about stitching together demos. It is about building systems that can reason over messy context, make good decisions, explain themselves, recover from failure, and improve over time. There is no separation here between product, engineering, and company building. You will help shape all three.
Concretely, you will be building the system itself — the runtime, the retrieval, the orchestration, the data infrastructure, the safety boundaries — not orchestrating prompts inside someone else's. Evaluation, observability, and quality loops are part of the work because you own the system end-to-end, not because someone hands you a pipeline to operate.
Build and ship agentic systems that reason across fragmented enterprise data and workflows
Design retrieval, memory, and context systems that help the product make better decisions over time
Build the infrastructure around the model: orchestration, safeguards, fallbacks, logging, replayability, and monitoring
Build the evaluation infrastructure that measures quality, regressions, trust, and real-world performance — and the feedback loops that use it to improve the system
Turn inconsistent signals, business logic, and historical outcomes into durable system context
Design systems that know when to act, when to ask, when to wait, and when to do nothing
Improve reliability under real production constraints like latency, cost, partial failures, and changing data
Work across the stack when needed, including backend systems, product surfaces, internal tools, and data flows
Partner closely with the founders on product direction, technical architecture, and engineering culture
Have 5+ years building production systems in one or more of: distributed backends, data pipelines, retrieval or agent runtimes, or core ML infrastructure
Have personally owned systems end-to-end — through design, ship, real failure modes, and recovery — that real users relied on
Are comfortable moving across backend, data, and AI application layers
Care deeply about correctness, edge cases, failure modes, and decision quality
Like ambiguous problems where the right abstraction is not obvious at the start
Move fast while keeping quality high
Want real ownership in a small, intense, high-trust team
Have strong product instincts and like building things people actually use
Strong signal: You've designed and shipped LLM systems end-to-end — the agents, retrieval, orchestration, and the evals you built to keep them honest — and you understand their failure modes from operating them in production, not from running someone else's benchmark.
You want a narrowly scoped role with clean boundaries
You need complete specs before you can start
You mainly want to do research rather than ship product
You are excited by demos more than reliability
You prefer polished environments over messy, early-stage ones
You optimize for local elegance over end-to-end outcomes
You will be one of the first engineers helping define the company and the product
You will work on problems where the answers do not exist yet
You will build systems that matter inside revenue-critical enterprise workflows
You will have real ownership across product, architecture, and execution
You will be working directly with an experienced founding team
You will help shape how autonomous decision systems are actually built in practice
Remote
Competitive salary and meaningful equity. All applicants must be authorized to work in the United States for any employer. Augmeta cannot sponsor work visas or permits.