SAP to open-stack ERP migration
Manufacturer, 500+ employees. Phased migration off SAP ERP onto an open stack with custom modules — run in stages while the plant kept working.
- 0 downtime
- 8 mo migration
- −40% on licences
The demo was fine. Then somebody pasted something clever into the chat, the model answered from another tenant's document, and the monthly bill quadrupled. Three failures, three different tests.
What we can show you
Injection, leakage, ungrounded answers, retrieval permissions, runaway cost and logging. Each has its own test and its own fix.
Six areas. Prompt injection and jailbreaks — can a user make the feature ignore its instructions or reveal them. Data leakage — can one customer see another's content through retrieval or history. Grounding — how often the answer is confidently wrong on your own questions. Permissions inside retrieval — does the index respect who may read what. Cost — what a hostile or careless user can spend in an hour. Logging and privacy — what is stored, for how long and with which personal data still in it.
This is not a penetration test of your application: that covers the perimeter and the code. This covers the model layer, where the failures are new and most checklists have nothing to say.
When an AI feature is already live without an evaluation suite; when answers are «sometimes odd» and nobody measures how often; when the model bill grows faster than usage. If the feature is still a prototype, fixing the architecture is cheaper than auditing it — see prototype rescue.

Price
$1,790
Timeline
5 working days
Payment
Credited against the fixes
They overlap in almost nothing, and most teams need both — usually in this order.
The deliverable that matters is not the report — it is the harness that keeps the report true after your next prompt change.
Prompt injection and jailbreak attempts against the live behaviour, including indirect injection through documents the model retrieves.
Retrieval permissions: we try to reach content the test user must not see, across tenants and across roles.
A set of questions from your own domain with known answers, run repeatedly, to measure how often the feature is confidently wrong.
Cost testing: worst-case tokens per request, what an abusive user can spend in an hour, and where a cap belongs in code.
Every finding has steps, an example, a severity and a proposed fix with an effort estimate. Nothing is reported that we could not reproduce twice.
Findings that are architectural rather than tactical are named as such — sometimes the honest answer is that the feature needs rebuilding, not patching.
The evaluation suite runs on your data in your CI. After any prompt, model or retrieval change you can re-run it and see what moved.
The method behind our measurements is public in the engineering report — including its limits.
Four things we can prove, rather than four adjectives.

Evaluation harnesses and guardrails are part of every AI project we ship. The audit is that discipline pointed at someone else's system.

Most AI security companies sell a one-off report, and a report ages the moment someone edits a prompt. A re-runnable suite on your data does not — that is the actual deliverable.

We do not publish «reduces hallucinations by N%» without a case behind it. Findings come with reproduction steps and a method you can check.

The audit fee is credited against the fixes, and the fixes are priced separately so you can take the report and fix it yourself.
Industries
The unacceptable failure differs by domain: a wrong quote in logistics, a leaked contract in property, a confidently wrong answer in customer support.
Six answers that usually replace a discovery call.
No. A pentest covers the application and infrastructure; this covers the model layer — injection, leakage through retrieval, grounding, permissions and cost. Most teams need both.
A report with reproducible cases, severity and proposed fixes — plus an evaluation harness that runs on your data and that you keep.
Usually not: a staging copy with representative data is enough, and it keeps customer data out of the testing loop.
Yes, quoted separately, with the audit fee credited against it. You are equally free to take the report and fix it in-house.
Any hosted or self-hosted LLM. The tests are about the system you built around the model, not about the vendor's benchmark scores.
Five working days for the audit, with the report and the harness handed over on the fifth.
Manufacturer, 500+ employees. Phased migration off SAP ERP onto an open stack with custom modules — run in stages while the plant kept working.
Shipped an AI feature without evaluations?
Tell us what the feature does and where it runs. You get the audit scope, what we will try to break, and a fixed price — the fee is credited against the fixes.
Tell us what you need — we'll suggest the format and the timeline. We normally reply within one business day.
or write to us directly