Anthropic and OpenAI just committed to embedded third-party evaluators →Remote-first · NDA before any engagement starts
Embedded AI safety & security evaluators

The embedded evaluator model Anthropic and OpenAI just committed to: we already run it.

Embedded Evals audits your models and agents for jailbreaks, prompt injection, and unsafe tool use, then stays embedded with the same kind of ongoing, employee-level access frontier labs just pledged to give independent evaluators, for companies who aren't frontier-lab scale but face the same pressure.

Model & framework agnostic
First finding in 5 business days
EU AI Act & NIST AI RMF mapped
eval-monitor · live
128Attacks blocked today
99.4%Guardrail coverage
3Open findings
EVL-0091Critical
EVL-0084High
EVL-0077Medium

Eval Monitor (illustrative)

The shift, dated

September 2026: embedded evaluation went from advisory to expected

On September 12, 2026, Anthropic CEO Dario Amodei published “We Must Pace the Frontier”, a three-part plan for slowing AI capability advances. Step one, which Anthropic is adopting unilaterally and immediately, is giving independent third-party evaluators ongoing, employee-level access (office space, badges, company laptops, tool permissions) to verify safety practices, report incidents, and assess model alignment and training pipelines, with the right to publish findings the company cannot veto.

“Employee-level access”: the phrase Amodei used for what a real third-party evaluator relationship requires. Not a report you commission. Ongoing access you grant.

OpenAI's Sam Altman publicly agreed the same day and committed his own company to the same kind of access. Coverage here.

Neither company framed this as something only frontier labs need. It's being set up as the standard the rest of the industry gets measured against. Embedded Evals runs that same model: an audit, then ongoing embedded access, for companies that aren't training frontier models but are already being asked who verifies theirs is safe.

To be direct: we are not affiliated with, endorsed by, or the same organization as Anthropic, OpenAI, or METR (the evaluator org Amodei names as an example of this role). We're an independent firm built around the same embedded-evaluator model, for companies that need it now and aren't Anthropic-scale.

5Business days to first finding
24/7Continuous monitoring on retainer
6 wkTypical pilot build
0Client data retained after engagement
Who this is for

Three kinds of teams end up needing this at the same moment

Right before something ships.

AI labs

Training or fine-tuning a model for real users? We run adversarial evaluations against the release candidate before your users find the gaps for you.

Agent builders

Your agent can call tools, touch customer data, or execute code. We test what it actually does when someone tells it to do something it shouldn't.

Regulated enterprises

Rolling out AI under the EU AI Act or sector guidance? We map your controls to the evidence an auditor or regulator will actually ask for.

What an engagement finds

A sample findings board

Illustrative findings from a typical six-week pilot, anonymised. Every stack is different, and yours may look nothing like this.

EVL-0091Found

Instructions embedded in a fetched webpage's alt text overrode the agent's system prompt and triggered an unauthorised refund tool call.

Critical
EVL-0084Guardrail shipped

End users could edit the internal search tool's description field, injecting instructions into every downstream call using it.

High
EVL-0102Guardrail shipped

Agent followed instructions embedded in a customer-uploaded PDF and attempted to email account data to an external address.

High
EVL-0058Verified closed

A crafted support ticket could impersonate an admin role in the tool-routing prompt, unlocking account-deletion actions.

High
EVL-0077Found

Model disclosed system-prompt contents under a multi-turn role-play framing in 6 of 20 adversarial attempts.

Medium
Where we work

Five surfaces, one attack library

Adversarial red-team assessment

Structured attacks (prompt injection, jailbreaks, tool misuse, data exfiltration, multi-turn manipulation), scored by what actually gets through.

Guardrail & policy engineering

Your rules turned into enforceable, testable controls: input/output filters, tool-permission boundaries, approval gates for anything irreversible.

Agentic action auditing

Every autonomous action gets logged and anomalies flagged, so you can answer "what did it actually do," not just "what did we tell it to do."

Continuous regression testing

Every model, prompt, or agent update is re-run against your attack library before it ships.

Compliance readiness

Controls mapped to the EU AI Act, NIST AI RMF, and sector guidance, with evidence a regulator will actually ask for.

How an engagement runs

Audit, then pilot, then embedded

1
1–2 weeks

Readiness Audit

We map your attack surface and run a baseline red-team pass. You get a findings report and a risk score, whether or not you go further.

2
6 weeks

Pilot Build

We implement guardrails and monitoring for your highest-severity findings, stand up an audit dashboard, and re-test until they're closed.

3
Ongoing, cancel anytime

Embedded Retainer

The same ongoing, employee-level access model Anthropic and OpenAI just committed to for their own evaluators: every release re-tested, monitoring continuous, and we're on call when it matters.

Pricing

Published rates, no sales call required

Readiness Audit

$6,500 one-time

A baseline red-team pass across your models and agents, plus a findings report with a risk score.

Secure checkout via Stripe. Prefer to talk first? Book a call instead.

Embedded Retainer

$12,000 / month

The embedded-evaluator model, run for you: ongoing access to verify what changes, continuous red-teaming on every release, and incident response on call. Cancel anytime.

Talk to us
FAQ

Before you ask

Will this slow down our release cycle?

The readiness audit runs alongside your existing schedule; nothing blocks a release unless you ask us to gate one. Continuous regression testing on the retainer typically adds under an hour per release.

Do you need our model weights or training data?

No. Red-teaming and guardrail testing run against your model's inputs and outputs, the same access a real attacker would have.

What if you find something critical mid-audit?

We tell you immediately, not in the final report. Critical findings don't wait for a deliverable date.

What AI stacks do you support?

Any model provider (OpenAI, Anthropic, open-weight, self-hosted) and any agent framework. If it takes instructions and can take actions, we can test it.

Can this replace our security team?

No, it extends it. We specialise in the attack surface that's specific to models and agents, which most existing security tooling wasn't built to see.

Is our data safe with you?

We work against a staging environment wherever possible, sign an NDA before any audit begins, and don't retain your data past the engagement. See our privacy policy for what we collect through this site.

Are you affiliated with Anthropic, OpenAI, or METR?

No. We reference Anthropic's and OpenAI's September 2026 commitments to embedded third-party evaluators because it's the reason this category now matters, not because we're connected to either company or to METR (the evaluator org named in that plan). Embedded Evals is an independent firm running the same model for companies that aren't frontier-lab scale.

Get started

Book the audit

Tell us where you are. We reply within one business day.

Prefer email? Write to us directly at info@voltrunnerev.com.