Adversarial capability evaluation

We find the work frontier models still get wrong

Capability red teaming: operational tasks built specifically to defeat a frontier model, where an expert in that field completes the same task with precision. The gap between those two outcomes is the finding, and it is the part that does not show up on a public benchmark.

Active in the micro1 Company Data Partnerships Program.

How we break a model

Failure is easy to produce and hard to produce usefully. A task that fails because it is badly specified teaches nobody anything. These fail because the work is genuinely difficult.

Long horizons and branching decisions

Real operational work does not resolve in one turn. It runs across many steps, changes direction on what it finds, and requires holding earlier context that has since gone stale. Models that look strong on single turn tasks come apart here, and that is where we aim.

Ambiguity that a practitioner resolves silently

Experts constantly fill gaps the brief never mentions, using judgement they would struggle to write down. We build tasks around exactly those gaps, because they are where a confident, plausible, wrong answer is most likely and most expensive.

A human ceiling on every task

Each task ships with a completion by someone who does that work professionally. That is what turns a failure into a finding: not that the task was hard, but that it was demonstrably doable and the model still did not do it.

What never leaves our side of the table

This work contributes operational expertise, not data about the people we work for. These are operating rules, applied before a workflow is written rather than checked afterwards.

Scrubbed or fictionalised from the start

Where a real workflow touches anything sensitive, it is rebuilt as a scrubbed or fictionalised version before any work begins. Not redacted afterwards.

Never client or customer material

No client, customer or employer confidential information. No proprietary material. Nothing covered by an NDA. This is the same commitment that governs our software delivery work.

No personal data, no credentials

No private or personal information about anyone, and no passwords, API keys or credentials in any prompt, file or environment.

Only tools we can properly grant

We work only in tools where the operator personally holds access and can safely extend that access to an AI system. Where a real process would run through a company system, we substitute a personal or fictionalised equivalent.

Never on unauthorised devices

Agent work does not run on any device we are not authorised to run it on.

The functions we can cover

Breadth from running these functions as a working software business, not from staffing up for a programme. We are also among the top one percent of official FlutterFlow partners globally, which is what betting early on a shift looks like when it pays off.

Software engineeringQA and testProduct and project deliveryCustomer supportSales and CRM operationsInternal operations

How an engagement runs

1

Scope

We agree the functions, the difficulty bar and the rubric dimensions with you before anything is built.

2

Build

Workflows are authored by the people who do the work, then scrubbed or fictionalised and checked against the data rules above.

3

Run

Each workflow is executed across the model set under matched conditions, with the expert baseline captured alongside.

4

Score and compare

Every run is graded on the agreed dimensions, then models are ranked against each other so the per dimension picture is legible.

5

Deliver

You receive the workflows, the runs, the scores and the comparison, in the structure your pipeline expects.

Questions partners ask

Is this safety red teaming or jailbreak testing?

Neither. This is capability red teaming: finding operational work that frontier models cannot yet do reliably. We do not test for harmful outputs, alignment failures, guardrail bypasses or prompt injection resistance, and we would point you elsewhere for that.

What counts as a successful adversarial task?

One an expert completes with precision and current frontier models do not. If everything passes it, the task is too easy to be informative. If nothing can pass it, including our own practitioner, it is badly specified rather than hard, and it does not ship.

Do you report the failures, or just the scores?

Both. A score tells you a model lost; the run tells you where and how, which is the part that changes what you train on next. Every task ships with its runs, the expert baseline and the per dimension scoring.

Who does the work?

The same practitioners who do the job for clients. The person who defines a QA workflow is a QA practitioner, and the person grading a software engineering run is an engineer. That is the point: the expert baseline has to be genuinely expert or the comparison is not worth much.

Do you ever use real client data?

No. Client confidentiality and intellectual property are non negotiable, and they do not bend for this programme. Where a workflow is drawn from real operational experience, it is rebuilt as a scrubbed or fictionalised version before the work starts, not cleaned up afterwards.

How do we start?

A short call to scope one batch. We would rather prove the work on a small, well defined set than talk about volume before you have seen the quality.

Related

Send us a function you think is already solved

The most useful first batch is usually the work a team assumes models have covered. Tell us the function and we will build the tasks that test it.

Start a conversation