DeepBlocker / Products / 05 · Trustwright
RED TEAM · PRODUCT 05

Trustwright

AI assistants now take actions on websites on your staff's behalf. Trustwright tests whether a dishonest website can talk them into the wrong one.

In one sentence
We try to socially engineer your AI assistants the way a malicious website would, then show you which manipulations worked and what they would have cost you.
What it is
A testa controlled attack, not software you install
Who it is for
Any organisationwhose staff use AI assistants in business systems
Target
Your AI agentsnot your people
Status
Live nowtrustwright.deepblocker.ai · open source

What just changed

Until this year, an AI assistant looking at a website had to squint at the screen and guess where to click. That has changed, and almost nobody has noticed the security consequence.

01

Websites now hand AI assistants a menu

A new web standard called WebMCP lets a website publish a list of actions it will let an AI assistant perform, written in plain English. Search products. Add payee. Approve payment. The assistant picks from the list instead of guessing.

02

It is real infrastructure, not a prototype

It is a W3C standard. Google Chrome ships it with a public origin trial for production sites, and ChatGPT's in-app browser supports it out of the box.

03

The assistant has to take the website at its word

Nothing verifies that an action does what its description says. A description can carry hidden instructions aimed at the assistant. A site can swap one action for another after the assistant has already committed to it.

04

And the assistant is holding your staff's authority

It works inside the logged-in session. A manipulated assistant is not a wrong answer on a screen. It is an action taken on a real account.

This is not a prediction. Google's own developer documentation names the attack classes and states they cannot be fully prevented. Independent academic testing published in 2026 manipulated the current generation of assistants, including models from OpenAI, Anthropic and Google, with success rates reaching 100 per cent for some techniques. The standards body has publicly asked the industry to build a shared test set for exactly this. Nobody has.

How the assessment works

A short, contained engagement. Nothing is installed on your systems and no real account is ever touched. We agree the scope, run the manipulations, and debrief.

01

Scope and consent

We agree which assistants are in scope, which business workflows matter most, and what a bad outcome would look like for you. Signed rules of engagement before anything runs.

02

We run the manipulation library

Your assistants are exposed to realistic dishonest websites drawn from the published attack classes: hidden instructions inside action descriptions, instructions buried in returned data, actions swapped after the assistant commits, and actions that lie about whether they change anything.

03

We measure, we do not guess

Every attempt is scored on what the assistant actually did, not on what it said. Either it took the action it should not have, or it did not. There is no interpretation in the result.

04

Debrief and evidence pack

You get a scored report, the business consequence of each successful manipulation, the specific changes that close it, and a sealed record you can hand to your board, your auditor or your insurer.

Every attempt is inert. Nothing is stolen, nothing is transmitted anywhere, and no live system is touched. We use harmless marker tokens, so we can prove precisely whether an assistant was manipulated while carrying no risk to you at all.

What you get

The same deliverable our voice-fraud clients already buy, in the format their auditors already accept.

01
Scored manipulation report
Every attempt, whether it succeeded, and what your assistant did in response.
02
Business consequence mapping
What each successful manipulation would have cost you in the workflow it was aimed at.
03
Remediation plan
The specific policy, configuration and control changes that close each finding.
04
Sealed evidence pack
Timestamped and hash-sealed, so you can prove the test happened and what it found.
05
Re-test on a schedule
Assistants change and so do the attacks. Quarterly re-testing shows the trend line.

Who this is for

Risk and compliance

Regulated financial firms

Trust companies, fund administrators, private banks and wealth managers whose staff are already running AI assistants inside client and payment systems.

Trigger: nobody can answer "have we tested that?"
Security

CISOs and heads of security

Teams that already test phishing and voice, and now have a third population of targets that never gets tired and never gets suspicious.

Trigger: AI assistant rollout has outrun the control review
Product and engineering

Teams publishing agent actions

If you are exposing actions to AI assistants on your own website, you need to know how assistants behave against a dishonest version of it before your customers find out.

Trigger: shipping agent support this quarter

Pricing

Fixed scope, delivered in two to three weeks. Early-access pricing is set with our first design partners.

Assessment
£9.5k-15k
one-off, fixed scope
  • Scoping workshop and rules of engagement
  • Full manipulation library run against your assistants
  • Scored report and business consequence mapping
  • Remediation plan and debrief
  • Sealed evidence pack
Continuous
£2.5k-4k
per quarter
  • Quarterly re-test as assistants and attacks change
  • Versioned evidence pack with trend line
  • New attack classes added as they are published
Start here

Find out whether your AI assistants can be talked into the wrong thing.

We are taking a small number of design partners now, at reference pricing, in exchange for a case study.