Founding Engineer, Agent Systems
London · Full-time
About HelmGuard
We're building governance and security infrastructure for a world of increasingly capable and widely proliferated agents.
The world is undergoing a fundamental shift: every day more work is being done autonomously by AI. This is true for enterprises but also for the adversarial actors who seek to disrupt their operations through cyber attacks. The only way to keep up is to use AI intelligently to govern internal deployments and defend against offensive threats. HelmGuard will be the platform which businesses use to do both.
We've grown to seven-figure revenue within 9 months of product launch, on the back of multi-year contracts with leading enterprises in financial services, regulated technology, and healthcare. Our founders come from Palantir and academic institutions: Oxford, Stanford, and ETH. We're backed by leading UK and US institutional investors and exceptional angels from Meta, Isomorphic Labs, Palantir, SpaceXAI, and more.
We're hiring across founding-team roles for people who want outsize impact, the influence over direction and culture that comes only from joining this early, and pre-Series A equity upside.
Non-Negotiables
In-office 4 days a week (with flexibility for things like school drop-off)
Wanting to own outcomes and to take on the hours required to do so
Having an existing right to work in the UK (we do provide visa support for applicants already working in the country)
Your Impact
We already have one of the most robust and high-performance agent scaffolds in our vertical. You’ll help make it one of the best AI agent platforms being shipped today. In addition, you’ll help define the primitives that fast-growing startups and massive government organisations use to govern agents.
The Role
You own the backend for our agent platform. This means you are responsible for the cost, performance, and reliability of the agents we use. What’s involved in that can expand with your ambition and our progress as a company. On day one, it involves further improvements to our evals, continued integration of our AI security primitives, and iteration on our harness. In the future, it will also likely involve research under our CTO to further improve our agent delivery and help govern our customers’ agents.
Importantly, this role is mainly engineering, with some room for focused research. We expect candidates to have run agents in production and to be ready to build for production (not research) workloads.
What You Will Do
Agent scaffolding: tool use, context management, sandboxing, prompt-injection defence
Evals for security and GRC workflows
Reliability infrastructure: retries, fallbacks, circuit breakers, prompt versioning
What You Bring
Experience with backend engineering in TypeScript or a comparable language, with 1–2 solid projects involving production LLM features
Experience with agent frameworks, tool calling, and multi-step orchestration
Good taste in eval construction and grading
Strong systems thinking: async, queues, idempotency
Nice To Have
Comfortable reading and implementing research from AI venues like ICML, NeurIPS, TMLR or ICLR
Background in GRC or security
Background in AI security, safety or governance
This Role Probably Isn’t For You If
You haven’t deployed an AI agent consumed by someone other than yourself and iterated on it using evaluations as a scorecard
You wouldn’t be able to confidently answer relatively niche questions about Anthropic’s API
You aren’t able to name an inference provider outside of the major labs
You've never had to make an agent system cheaper
You’ve shipped an LLM feature but not built an eval for it
Culture and Values
We value a diversity of perspectives and experiences. We also hold a small set of core beliefs that reflect how we operate, and share them transparently with candidates so the fit is clear from the outset.
Put Customers First. Our customers buy outcomes from us, not features. We judge every decision by whether it delivers on that promise.
Take Ownership. Founding-stage means problems don't come pre-scoped. You see something that needs doing, scope it, ship it, own the outcome. We expect this from everyone, and provide you with the backing to execute on it.
Work Hard, With Gratitude. This is the most consequential window ever for building a business. We work hard because the opportunity is rare, and we do it with gratitude for the moment, for the people we get to build it with, and for the customers willing to bet on us this early.
Say the Silly Thing. The best ideas usually start out sounding half-baked, so we'd rather you say the silly thing than sit on it. We want you opinionated and willing to argue, and just as willing to change your mind when someone makes a better case. Disagreement here is a contribution, not a risk.
Working at HelmGuard
Location. King's Cross, London (Gridiron building). We put a large value on in-person collaboration. Our default policy is 4 days a week in-office, with one day remote (i.e. 4+1).
Compensation. Top decile for the London market, with meaningful EMI-eligible options.
Perks. Daily team lunch and specialty coffee, a roof terrace overlooking King's Cross, on-site showers for those who enjoy active commuting, and serious per-engineer AI tooling and API budgets.
Interview process. Three stages: behavioural phone screen, technical phone screen, and a paid on-site work trial. Target turnaround is under two weeks from first conversation.
Tech Stack. TypeScript, Node.js, React, Tailwind, OpenAPI, Express, Azure (Container Apps, Service Bus, Front Door, Entra ID), Postgres, Terraform, GitHub Actions, Docker. Anthropic-first AI with in-house evals and scaffolding. Claude Code throughout, but we also use Codex and are open to other scaffolds.
Relocation. We are not currently in a position to help candidates relocate to London from outside the country if they require visa sponsorship to do so. We can help candidates who already hold a UK visa transfer it to us.