Adversarial capability evaluation
We find the work frontier models still get wrong
Capability red teaming: operational tasks built specifically to defeat a frontier model, where an expert in that field completes the same task with precision. The gap between those two outcomes is the finding, and it is the part that does not show up on a public benchmark.
Active in the micro1 Company Data Partnerships Program.
How we break a model
Failure is easy to produce and hard to produce usefully. A task that fails because it is badly specified teaches nobody anything. These fail because the work is genuinely difficult.
Long horizons and branching decisions
Real operational work does not resolve in one turn. It runs across many steps, changes direction on what it finds, and requires holding earlier context that has since gone stale. Models that look strong on single turn tasks come apart here, and that is where we aim.
Ambiguity that a practitioner resolves silently
Experts constantly fill gaps the brief never mentions, using judgement they would struggle to write down. We build tasks around exactly those gaps, because they are where a confident, plausible, wrong answer is most likely and most expensive.
A human ceiling on every task
Each task ships with a completion by someone who does that work professionally. That is what turns a failure into a finding: not that the task was hard, but that it was demonstrably doable and the model still did not do it.
What never leaves our side of the table
This work contributes operational expertise, not data about the people we work for. These are operating rules, applied before a workflow is written rather than checked afterwards.
Scrubbed or fictionalised from the start
Where a real workflow touches anything sensitive, it is rebuilt as a scrubbed or fictionalised version before any work begins. Not redacted afterwards.
Never client or customer material
No client, customer or employer confidential information. No proprietary material. Nothing covered by an NDA. This is the same commitment that governs our software delivery work.
No personal data, no credentials
No private or personal information about anyone, and no passwords, API keys or credentials in any prompt, file or environment.
Only tools we can properly grant
We work only in tools where the operator personally holds access and can safely extend that access to an AI system. Where a real process would run through a company system, we substitute a personal or fictionalised equivalent.
Never on unauthorised devices
Agent work does not run on any device we are not authorised to run it on.
The functions we can cover
Breadth from running these functions as a working software business, not from staffing up for a programme. We are also among the top one percent of official FlutterFlow partners globally, which is what betting early on a shift looks like when it pays off.
How an engagement runs
Scope
We agree the functions, the difficulty bar and the rubric dimensions with you before anything is built.
Build
Workflows are authored by the people who do the work, then scrubbed or fictionalised and checked against the data rules above.
Run
Each workflow is executed across the model set under matched conditions, with the expert baseline captured alongside.
Score and compare
Every run is graded on the agreed dimensions, then models are ranked against each other so the per dimension picture is legible.
Deliver
You receive the workflows, the runs, the scores and the comparison, in the structure your pipeline expects.
Questions partners ask
Is this safety red teaming or jailbreak testing?
Neither. This is capability red teaming: finding operational work that frontier models cannot yet do reliably. We do not test for harmful outputs, alignment failures, guardrail bypasses or prompt injection resistance, and we would point you elsewhere for that.
What counts as a successful adversarial task?
One an expert completes with precision and current frontier models do not. If everything passes it, the task is too easy to be informative. If nothing can pass it, including our own practitioner, it is badly specified rather than hard, and it does not ship.
Do you report the failures, or just the scores?
Both. A score tells you a model lost; the run tells you where and how, which is the part that changes what you train on next. Every task ships with its runs, the expert baseline and the per dimension scoring.
Who does the work?
The same practitioners who do the job for clients. The person who defines a QA workflow is a QA practitioner, and the person grading a software engineering run is an engineer. That is the point: the expert baseline has to be genuinely expert or the comparison is not worth much.
Do you ever use real client data?
No. Client confidentiality and intellectual property are non negotiable, and they do not bend for this programme. Where a workflow is drawn from real operational experience, it is rebuilt as a scrubbed or fictionalised version before the work starts, not cleaned up afterwards.
How do we start?
A short call to scope one batch. We would rather prove the work on a small, well defined set than talk about volume before you have seen the quality.
Related
Send us a function you think is already solved
The most useful first batch is usually the work a team assumes models have covered. Tell us the function and we will build the tasks that test it.
Start a conversation


