External attack surface
Domains, certificates, DNS, ports, services and cloud assets reconciled, including shadow infrastructure and impersonation.
It begins with no credentials and no source code — the outsider's view — mapping what your estate exposes and testing what that reaches. Grant code, cloud and identity access and it goes deeper. Either way, a finding ships only with evidence behind it.
Every assessment is authorized, scoped and non-destructive. AATA proves risk without weaponising it, and reports findings rather than certifying a system. Consoles shown on this page are illustrations with sample data.
How much context you supply changes what can be proven and what is allowed to run. The first two modes need nothing from you but written permission.
Publicly observable data only. What your estate leaks to anyone looking, with no interaction beyond ordinary retrieval.
No credentials, no code. Authorized testing of your external attack surface exactly as an outsider would reach it.
A constrained test account. Authenticated workflows, tenant boundaries and the paths a low-privilege user could take.
Code, cloud, identity and configuration. Deep design and implementation review through read-only connectors.
Code plus forked chain state. Invariant and economic simulation, where transactions execute only inside the fork.
Staging and synthetic data. Release assurance, where active testing is permitted within the agreed policy.
Event integrations on approved assets. Incremental checks on every material change, escalating when risk moves.
A time-bound incident scope. Evidence correlation and containment advice, human-led, with no autonomous action.
The same path whether it runs once before a launch or on every merge. The first two stages happen before an agent touches anything.
The Scope Guardian resolves a machine-readable manifest: legal owner, included and excluded assets, permitted testing classes, rate limits and windows.
The Asset Cartographer graphs domains, certificates, cloud accounts, identities and contracts. Change Intelligence isolates what moved since the last run.
The Mission Planner splits the work into parallel workstreams, each with an agent, a model, a tool budget and a deadline. Restricted data can be pinned local.
Specialists work their own paths across surface, code, cloud, contracts and money flows, publishing every observation to a shared evidence graph.
The Adversarial Challenger argues against each hypothesis. Contradictory evidence blocks a finding even when several agents agree on it.
Findings are scored, routed to Jira, Linear, ServiceNow or Slack with an owner, and re-tested on the deployment that claims to have fixed them.
Domains, certificates, DNS, ports, services and cloud assets reconciled, including shadow infrastructure and impersonation.
Authorized testing of web, API, GraphQL and mobile backends: authentication, tenant boundaries and business logic.
Static and semantic analysis across repositories, with cross-repo data flow, secrets exposure and release deltas.
AWS, Azure and GCP posture, Kubernetes, and IAM graph analysis that surfaces privilege-escalation paths.
Solidity, Vyper, Rust, Move and Cairo reviewed with symbolic execution, fuzzing and forked-chain simulation.
Hot, warm and cold architecture, MPC and HSM signing quorums, and the deposit, withdrawal and sweep flows.
Router, pool, LP, hook and solver review, including MEV exposure, pricing paths and slippage handling.
Oracle manipulability, flash-loan feasibility and liquidation paths modelled against real liquidity depth.
Double-entry invariants, reconciliation, idempotency, retries and race conditions across the whole lifecycle.
FIX, SWIFT, ISO 20022, ACH, RTP and open banking reviewed for integrity and entitlement enforcement.
Order entry, matching, margin, liquidation and settlement examined as one chain, not separate systems.
Rules, models, alerts and case flows exercised as live controls, to establish what actually gets through.
A merge, deployment, IAM change, contract upgrade or signer rotation triggers a targeted reassessment.
A build can be blocked on an unresolved critical finding, set per environment and per severity.
The original evidence is replayed against the new code before a finding is ever marked verified.
Proven attack paths are recomputed on every material change, so a reintroduction surfaces as a regression.
Each task routes to a specialist, and the specialists have to argue. Seven of the seventeen roles are shown here.
Holds a veto over every active action. Enforces the authorization manifest, exclusions, rate limits, testing windows and stop conditions, and writes every refusal to the audit ledger.
A withdrawal endpoint re-checks the daily limit before applying it, leaving a window where two concurrent requests both pass. Here is how that becomes a finding, and how it stops being one.
The run starts against a manifest, not a target list: who authorized it, what is in scope, what is explicitly out, and which classes of testing are permitted in each.

Agreement between agents is useful and insufficient. Every claim carries an evidence tier, and contradictory evidence can block one that all of them believe.
Three ways to look for the same vulnerability. A scanner is precise inside one domain and blind outside it. A single assistant reasons across domains with no mechanism to disprove itself.
| # | Capability | Conventional scanner | Single AI assistant | Autonomous AI Threat Assembly |
|---|---|---|---|---|
| 1 | Testing with no access at all | Surface scanning only | Needs to be told the target | Discovers, then validatesblack-box from one domain |
| 2 | Deterministic security tools | Strong in one domain | Often limited | Orchestrated across domainsversioned and sandboxed |
| 3 | Cross-domain reasoning | Low | Medium to high | Highgraph- and evidence-backed |
| 4 | Independent verification | Rule-dependent | Usually weak | Challenger and verifieras first-class agents |
| 5 | Business and financial context | Limited | Prompt-dependent | Structured and persistentcarried between runs |
| 6 | Smart-contract economic analysis | Specialist tools only | Inconsistent | Dedicated agents and forksinvariant and economic simulation |
| 7 | TradFi workflow expertise | Limited | Generic | Encoded invariantsledger, settlement, entitlements |
| 8 | Continuous change awareness | Scan-based | Usually ad hoc | Event-drivenincremental, risk-based |
| 9 | Privacy and model choice | Not applicable | Provider-dependent | Policy-based routinglocal, private or external |
| 10 | Reproducible evidence lineage | Good for tool output | Often weak | End to endclaim through to evidence |
| 11 | Enterprise workflow integration | Varies | Limited | Connector and event platformticketing, SIEM, SOAR, GRC |
| 12 | Authorization and safety policy | Product-specific | Often informal | Machine-enforcedScope Guardian holds a veto |
Inside a narrow domain a good scanner is fast, cheap and deterministic. AATA runs several as tools rather than competing with them. What they cannot do is carry business context across a boundary.
A frontier model reasons across domains well, and is confidently wrong at a rate no security programme can absorb. Without a challenger and deterministic corroboration, confidence is the only support.
Independent investigation paths, structured claims instead of a shared conversation, a challenger whose job is to disprove, and an evidence bar a finding must clear before anyone is paged.
Read-only scopes by default, least privilege per connector, short-lived tokens and a kill switch on every connection. These are supported integration targets across source control, cloud, identity, security operations, chains and payment rails.
Six questions that decide an autonomous testing evaluation, answered the way we would want them answered if it were our estate.
Built for organisations where a security failure moves money, breaks a regulated obligation, or ends a protocol.
Invariants, exploit paths and funds at risk before a launch or an upgrade, not after the governance vote.
Withdrawal flows, signing quorums, listing pipelines and reconciliation reviewed as one chain.
Ledger invariants, entitlements and settlement controls tested against the workflows a regulator asks about.
Code, cloud, identity and supply chain in one graph, plus prompt injection on the AI features you ship.
Continuous coverage between human engagements, so the next engagement starts from a known state.
Control mapping, coverage history and evidence integrity, so an assessment produces the audit artefact.
Autonomous assessment is one layer of a security programme. These are the pages for the controls it tests and the operations it feeds.
Autonomous AI penetration testing uses coordinated AI agents to discover, investigate and validate security weaknesses without a person driving each step. Unlike a scanner it reasons about how a system works and chains issues into attack paths. Unlike a single model, every claim must be supported by reproducible evidence before it becomes a finding.
Yes. Passive external and black-box modes need no credentials and no source code. AATA works from what your estate exposes publicly, discovers assets you may not know you own, and tests the authorized external surface the way an outsider would reach it. Code, cloud and identity access is optional and adds depth rather than being the starting point.
It is non-destructive by default. AATA proves risk using static evidence, read-only queries, synthetic canaries, isolated reproductions, chain forks and staging environments rather than exploitation. Destructive testing, persistence, denial of service and production data modification are prohibited actions enforced by a dedicated Scope Guardian agent, not left to agent judgement.
Every claim moves through a lifecycle: hypothesis, observed, corroborated, validated, accepted. An Adversarial Challenger argues against it and an Evidence Verifier replays or independently corroborates it. Findings carry an evidence tier from E0 to E4, and a critical designation requires E4, meaning multiple independent confirmations plus proof of impact.
A machine-readable rules-of-engagement manifest defines the legal owner, included and excluded assets, permitted testing classes, rate limits, testing windows and blackout periods. Agents run in sandboxes behind egress allowlists with short-lived scoped credentials, and a kill switch stops a run immediately. No active test starts before that manifest resolves.
A one-time engagement describes a system as it was during one particular week. AATA reassesses on the events that actually change risk, such as a merge, a deployment, an IAM change, a contract upgrade or a signer rotation, and it re-verifies fixes. Human expertise is not replaced; the Managed Security Program adds analyst validation on top.
AI can review contract code, storage layout, access control, upgrade paths and economic assumptions, and can run fuzzing, symbolic execution and fork simulation far faster than a person. It does not replace an accredited audit. AATA produces findings, reproduction evidence and remediation, and re-verifies fixes. It does not issue audit certificates.
That is a routing decision you control. Every model version in the registry declares its data classification limit, permitted regions and retention policy, and policy can pin secrets, keys, personal data or regulated content to local or customer-hosted models. AATA also runs in your VPC, fully self-hosted, or as a clean room destroyed after the mission.
It generates the evidence and the mapping: authorization history, assessment coverage, tool and model versions, finding lifecycle, remediation results and control drift. Mappings include SOC 2, ISO 27001, NIST CSF and 800-53, PCI DSS, DORA, NIS2 and NYDFS. Certification itself remains a decision for your auditor.
Bring your architect and whoever owns the authorization. We agree the estate, the exclusions and the safety policy first, then run a baseline and walk the findings with the evidence attached.
Scope an Assessment