The embedded evaluator model Anthropic and OpenAI just committed to: we already run it.
Embedded Evals audits your models and agents for jailbreaks, prompt injection, and unsafe tool use, then stays embedded with the same kind of ongoing, employee-level access frontier labs just pledged to give independent evaluators, for companies who aren't frontier-lab scale but face the same pressure.
Eval Monitor (illustrative)
September 2026: embedded evaluation went from advisory to expected
On September 12, 2026, Anthropic CEO Dario Amodei published “We Must Pace the Frontier”, a three-part plan for slowing AI capability advances. Step one, which Anthropic is adopting unilaterally and immediately, is giving independent third-party evaluators ongoing, employee-level access (office space, badges, company laptops, tool permissions) to verify safety practices, report incidents, and assess model alignment and training pipelines, with the right to publish findings the company cannot veto.
“Employee-level access”: the phrase Amodei used for what a real third-party evaluator relationship requires. Not a report you commission. Ongoing access you grant.
OpenAI's Sam Altman publicly agreed the same day and committed his own company to the same kind of access. Coverage here.
Neither company framed this as something only frontier labs need. It's being set up as the standard the rest of the industry gets measured against. Embedded Evals runs that same model: an audit, then ongoing embedded access, for companies that aren't training frontier models but are already being asked who verifies theirs is safe.
To be direct: we are not affiliated with, endorsed by, or the same organization as Anthropic, OpenAI, or METR (the evaluator org Amodei names as an example of this role). We're an independent firm built around the same embedded-evaluator model, for companies that need it now and aren't Anthropic-scale.
Three kinds of teams end up needing this at the same moment
Right before something ships.
AI labs
Training or fine-tuning a model for real users? We run adversarial evaluations against the release candidate before your users find the gaps for you.
Agent builders
Your agent can call tools, touch customer data, or execute code. We test what it actually does when someone tells it to do something it shouldn't.
Regulated enterprises
Rolling out AI under the EU AI Act or sector guidance? We map your controls to the evidence an auditor or regulator will actually ask for.
A sample findings board
Illustrative findings from a typical six-week pilot, anonymised. Every stack is different, and yours may look nothing like this.
Instructions embedded in a fetched webpage's alt text overrode the agent's system prompt and triggered an unauthorised refund tool call.
CriticalEnd users could edit the internal search tool's description field, injecting instructions into every downstream call using it.
HighAgent followed instructions embedded in a customer-uploaded PDF and attempted to email account data to an external address.
HighA crafted support ticket could impersonate an admin role in the tool-routing prompt, unlocking account-deletion actions.
HighModel disclosed system-prompt contents under a multi-turn role-play framing in 6 of 20 adversarial attempts.
MediumFive surfaces, one attack library
Adversarial red-team assessment
Structured attacks (prompt injection, jailbreaks, tool misuse, data exfiltration, multi-turn manipulation), scored by what actually gets through.
Guardrail & policy engineering
Your rules turned into enforceable, testable controls: input/output filters, tool-permission boundaries, approval gates for anything irreversible.
Agentic action auditing
Every autonomous action gets logged and anomalies flagged, so you can answer "what did it actually do," not just "what did we tell it to do."
Continuous regression testing
Every model, prompt, or agent update is re-run against your attack library before it ships.
Compliance readiness
Controls mapped to the EU AI Act, NIST AI RMF, and sector guidance, with evidence a regulator will actually ask for.
Audit, then pilot, then embedded
Readiness Audit
We map your attack surface and run a baseline red-team pass. You get a findings report and a risk score, whether or not you go further.
Pilot Build
We implement guardrails and monitoring for your highest-severity findings, stand up an audit dashboard, and re-test until they're closed.
Embedded Retainer
The same ongoing, employee-level access model Anthropic and OpenAI just committed to for their own evaluators: every release re-tested, monitoring continuous, and we're on call when it matters.
Published rates, no sales call required
Readiness Audit
A baseline red-team pass across your models and agents, plus a findings report with a risk score.
Secure checkout via Stripe. Prefer to talk first? Book a call instead.
Pilot Build
Guardrails and monitoring implemented for your highest-severity findings, re-tested until closed. Includes the audit.
Get startedEmbedded Retainer
The embedded-evaluator model, run for you: ongoing access to verify what changes, continuous red-teaming on every release, and incident response on call. Cancel anytime.
Talk to usBefore you ask
Will this slow down our release cycle?
The readiness audit runs alongside your existing schedule; nothing blocks a release unless you ask us to gate one. Continuous regression testing on the retainer typically adds under an hour per release.
Do you need our model weights or training data?
No. Red-teaming and guardrail testing run against your model's inputs and outputs, the same access a real attacker would have.
What if you find something critical mid-audit?
We tell you immediately, not in the final report. Critical findings don't wait for a deliverable date.
What AI stacks do you support?
Any model provider (OpenAI, Anthropic, open-weight, self-hosted) and any agent framework. If it takes instructions and can take actions, we can test it.
Can this replace our security team?
No, it extends it. We specialise in the attack surface that's specific to models and agents, which most existing security tooling wasn't built to see.
Is our data safe with you?
We work against a staging environment wherever possible, sign an NDA before any audit begins, and don't retain your data past the engagement. See our privacy policy for what we collect through this site.
Are you affiliated with Anthropic, OpenAI, or METR?
No. We reference Anthropic's and OpenAI's September 2026 commitments to embedded third-party evaluators because it's the reason this category now matters, not because we're connected to either company or to METR (the evaluator org named in that plan). Embedded Evals is an independent firm running the same model for companies that aren't frontier-lab scale.
Book the audit
Tell us where you are. We reply within one business day.
Prefer email? Write to us directly at info@voltrunnerev.com.