AI security audit for features that are already in front of customers

The demo was fine. Then somebody pasted something clever into the chat, the model answered from another tenant's document, and the monthly bill quadrupled. Three failures, three different tests.

AI audit

Six failure modes, one report you can act on

Injection, leakage, ungrounded answers, retrieval permissions, runaway cost and logging. Each has its own test and its own fix.

What an AI security audit actually checks

Six areas. Prompt injection and jailbreaks — can a user make the feature ignore its instructions or reveal them. Data leakage — can one customer see another's content through retrieval or history. Grounding — how often the answer is confidently wrong on your own questions. Permissions inside retrieval — does the index respect who may read what. Cost — what a hostile or careless user can spend in an hour. Logging and privacy — what is stored, for how long and with which personal data still in it.

This is not a penetration test of your application: that covers the perimeter and the code. This covers the model layer, where the failures are new and most checklists have nothing to say.

When an AI audit is worth doing

When an AI feature is already live without an evaluation suite; when answers are «sometimes odd» and nobody measures how often; when the model bill grows faster than usage. If the feature is still a prototype, fixing the architecture is cheaper than auditing it — see prototype rescue.

Signals your AI feature needs an audit

  • the feature is in production and there are no evals at all;
  • request logs are stored with customer data still in them;
  • answers are «sometimes strange» and nobody counts how often;
  • the model bill is growing faster than usage;
  • there is no rate limit or length cap on user input.
Signals an AI feature needs an audit: no evaluations in production, raw logs with customer data, unmeasured wrong answers and a growing model bill

Price

$1,790

Timeline

5 working days

Payment

Credited against the fixes

Pricing

What an AI audit costs with us

One fixed price for the audit, credited in full against the fixes. Model and inference costs during testing are passed through at cost.

Fixes

$7,990

Guardrails, retrieval permissions, grounding improvements, cost caps and rate limits, with the harness proving the change

From 2 weeks

AI feature rebuild

$10,900

When the architecture is the problem: retrieval, prompts and evaluation rebuilt around your data

From 3 weeks

Continuous evaluation

$1,890/month

The harness run on every change, drift and cost reports, a monthly review

Monthly

Hourly work $28–55 by role. See the full price list.

Comparison

Application penetration test or AI audit

They overlap in almost nothing, and most teams need both — usually in this order.

Criterion
PentestAI security audit
AI auditClassic penetration test
What is tested
Perimeter, application, infrastructure
The model layer: prompts, retrieval, guardrails
Prompt injection and jailbreaks
Rarely covered
Core of the work
Cross-tenant leakage through retrieval
Out of scope
Explicitly tested
Hallucination rate on your questions
Not a security concept there
Measured
Cost abuse
Not covered
Tested and capped
What you keep afterwards
A report
A report plus a re-runnable harness
Process

Five days, and what you have on day six

The deliverable that matters is not the report — it is the harness that keeps the report true after your next prompt change.

Days 1–2: attack surface

Prompt injection and jailbreak attempts against the live behaviour, including indirect injection through documents the model retrieves.

Retrieval permissions: we try to reach content the test user must not see, across tenants and across roles.

Days 3–4: grounding and cost

A set of questions from your own domain with known answers, run repeatedly, to measure how often the feature is confidently wrong.

Cost testing: worst-case tokens per request, what an abusive user can spend in an hour, and where a cap belongs in code.

Day 5: report with reproducible cases

Every finding has steps, an example, a severity and a proposed fix with an effort estimate. Nothing is reported that we could not reproduce twice.

Findings that are architectural rather than tactical are named as such — sometimes the honest answer is that the feature needs rebuilding, not patching.

What you keep: the harness

The evaluation suite runs on your data in your CI. After any prompt, model or retrieval change you can re-run it and see what moved.

The method behind our measurements is public in the engineering report — including its limits.

Why us

Why this audit and not a checklist

Four things we can prove, rather than four adjectives.

We build these features ourselves

Evaluation harnesses and guardrails are part of every AI project we ship. The audit is that discipline pointed at someone else's system.

You keep the harness

Most AI security companies sell a one-off report, and a report ages the moment someone edits a prompt. A re-runnable suite on your data does not — that is the actual deliverable.

No scare-percentage marketing

We do not publish «reduces hallucinations by N%» without a case behind it. Findings come with reproduction steps and a method you can check.

Fixes quoted, not implied

The audit fee is credited against the fixes, and the fixes are priced separately so you can take the report and fix it yourself.

Industries

AI auditby industry

The unacceptable failure differs by domain: a wrong quote in logistics, a leaked contract in property, a confidently wrong answer in customer support.

18 sectors

  • IT companies and digital agencies
  • Insurers and brokers
  • Lenders and microfinance
  • E-commerce
  • Wholesale and distribution
  • Manufacturing
  • Freight and forwarding
  • Clinics and medical centres
  • Dental
  • Car dealers and service centres
  • EdTech
  • Legal and consulting firms
  • Equipment rental
  • Real estate agencies
  • Property developers
  • Travel and events
  • Beauty and wellness
  • Construction and renovation
FAQ

AI audit questions

Six answers that usually replace a discovery call.

Is this a penetration test?

No. A pentest covers the application and infrastructure; this covers the model layer — injection, leakage through retrieval, grounding, permissions and cost. Most teams need both.

What do we actually receive?

A report with reproducible cases, severity and proposed fixes — plus an evaluation harness that runs on your data and that you keep.

Do you need production access?

Usually not: a staging copy with representative data is enough, and it keeps customer data out of the testing loop.

Can you fix what you find?

Yes, quoted separately, with the audit fee credited against it. You are equally free to take the report and fix it in-house.

Which models do you cover?

Any hosted or self-hosted LLM. The tests are about the system you built around the model, not about the vendor's benchmark scores.

How long does it take?

Five working days for the audit, with the report and the harness handed over on the fifth.

2–4×
faster on the code-writing share
in 3 weeks
from idea to MVP
230+
projects delivered
17 years
of engineering experience
Work

Work we can show you

Shipped an AI feature without evaluations?

Tell us what the feature does and where it runs. You get the audit scope, what we will try to break, and a fixed price — the fee is credited against the fixes.