ClearPenTest

What we test

AI and LLM Penetration Testing

Testing the AI surface of a product as a security surface: where untrusted content reaches a model, what the model is allowed to do with tools, and what a crafted input can make happen.

Who asks for this

Companies shipping AI features who are now being asked about them in security questionnaires, and companies whose agents can take actions rather than only produce text.

What the test covers

  • Indirect prompt injection through documents, web content, email and tool output
  • Tool and function authority: what the model can call, with whose permissions
  • Data boundaries between tenants when context, embeddings or caches are shared
  • System prompt and context exposure
  • Output handling, where model output reaches a browser, a shell, a query or another system
  • Agent action paths where untrusted input can reach a privileged or irreversible operation

What this test typically finds

Classes of finding, not a severity table. These are the issues that recur on this surface.

Untrusted content becoming instruction

Text from a document, page or API response that the model follows as a command. The core AI vulnerability, and the one that turns a chat feature into an action risk.

Tools with more authority than the user

An agent calling a backend with a service credential rather than the caller's, so the model can reach data the user could not.

Context bleed between tenants

Retrieval or caching that lets one customer's content appear in another customer's session.

Model output trusted downstream

Generated text inserted into a page, a query or a command without treating it as untrusted input, which is where prompt injection becomes code execution.

Questions people ask

What is AI penetration testing?

Security testing of the AI surface of an application: where untrusted content can reach a model, what tools the model can invoke and with whose permissions, whether context leaks between customers, and whether model output is trusted by downstream systems. It is a scope extension to an application test, not a separate discipline.

Is prompt injection a real vulnerability or a research problem?

It is a real vulnerability wherever a model has authority to act. Published vendor red-team work, including OpenAI's Operator system card, documents measurable susceptibility to indirect prompt injection in agents that use tools. The risk is the path from attacker-controlled instruction to a sensitive action, which is an ordinary security problem with a new entry point.

Does SOC 2 or ISO 27001 cover AI features?

Neither names AI specifically, but if an AI feature is part of the system in scope then it is part of what has to be tested. Enterprise security questionnaires have moved faster than the frameworks and now ask about model access, data use and tenant isolation directly.

Get a scoped price without a discovery call

Tell us what is in scope and what your audit needs. You get a fixed price and a date, not a quote after two meetings.