We spent our careers proving safety-critical systems work before they shipped.
Marker’s team comes from autonomous vehicles and large-scale production infrastructure. In those industries, every release had to survive simulation, regression testing, and safety review before it touched the real world. We built the simulation harnesses, evaluation pipelines, and internal tooling those teams ran on.
AI agents are now shipping into the same kind of high-stakes environments without that discipline. Marker brings it: repeatable tests, failure evidence, and an auditable record for every agent you put in front of customers.
Join us
We’re a small team in San Francisco, building in person.
We are unreasonably interested in voice AI making it into the real world — not the demo on stage, but the agent that answers a real phone line at 2am, gets interrupted, mishears an address, and has to recover without a human watching. That gap between a good demo and a system you would put in front of your own customers is the entire problem we work on.
The team is small and the surface area is large, so everyone owns things end to end: the API, the interface, the tests, the docs. We would rather delete something and rebuild it than defend a workaround. We talk to the people running these agents in production every week, and what they tell us changes the roadmap.
There is no long list of open roles. If this is the problem you want to spend the next few years on, write to us and tell us what you have built and what you would want to own here.