CISO AI Governance

Why IRAP-Style Security Assessments Make More Sense for AI and Cloud Risk

HackWednesday AI Security Desk2026-08-23

CISO AI GovernanceAI-generated draftAwaiting editor review6 verified source(s)

IRAP is useful because it pushes security buyers past generic compliance badges and toward scoped, evidence-based, risk-informed assessment reports that explain what was tested, what remains risky, and who is accountable.

The HackWednesday owl standing in a forest of scoped security controls, representing evidence-based IRAP-style assessment.
The new security question is not: are they certified? It is: what was assessed, against which controls, inside which boundary, with what remaining risk?
Editorial note: This AI-assisted article is published without a completed human review and should be read with extra scrutiny.

Security buyers are tired of shiny badges. A vendor can say it is secure, enterprise-ready, AI-safe, compliant, trusted, hardened, and cloud-native, but none of those words answer the question a CISO actually has to defend: what was assessed, against which controls, inside which boundary, using what evidence, and what risk remains?

That is why IRAP-style security assessment thinking makes more sense for the next generation of AI, SaaS, and cloud security reviews. IRAP, the Infosec Registered Assessors Program run by the Australian Signals Directorate through the Australian Cyber Security Centre, gives organisations access to independent assessors who can evaluate systems, cloud services, gateways, and related environments against the Australian Government Information Security Manual and related guidance.

The important nuance is that IRAP is not magic fairy dust. ASD states that IRAP assessors do not accredit, certify, endorse, or register systems on ASD's behalf. A completed IRAP assessment does not automatically mean a system is compliant with every tested control. Customers still need to read the report, understand the scope, review the tested controls, and decide whether the residual risk fits their own environment.

That limitation is exactly why the model is useful. Modern security assessments should not pretend to produce universal truth. They should produce decision-grade evidence. A useful assessment tells the buyer what was in scope, what was out of scope, which services were reviewed, which control objectives were met, which were partially met, which were not met, what compensating controls exist, and what the customer remains responsible for under the shared responsibility model.

For AI and agentic systems, this matters even more. A traditional vendor questionnaire often asks whether a control exists. An IRAP-style assessment asks a better chain of questions: where does the model run, what data can it access, which identities can invoke tools, how are prompts and tool calls logged, what can the agent modify, how is egress restricted, how are secrets protected, how are model outputs reviewed, and how does the provider prove those answers?

The phrase 'IRAP score' should be used carefully. IRAP itself is not a single public score like a credit rating. But security teams can borrow the idea of a scorecard from IRAP-style evidence: scope quality, control coverage, evidence freshness, assessor independence, residual risk clarity, customer responsibility clarity, and remediation maturity. That is more useful than a binary badge because it helps buyers compare vendors without pretending every system has the same risk profile.

A practical IRAP-inspired vendor scorecard could use seven dimensions. First, assessment boundary: does the report clearly show which products, regions, services, support environments, and third parties were assessed? Second, control mapping: are controls mapped to a recognised framework such as the ISM, ISO 27001, SOC 2, NIST, or cloud-specific guidance? Third, evidence quality: are claims backed by screenshots, logs, architecture diagrams, policies, tests, and interviews rather than marketing language?

Fourth, freshness: how old is the evidence, and have there been major architecture, ownership, service, legal, region, or incident changes since the assessment? Fifth, shared responsibility: does the report explain which controls belong to the provider and which controls belong to the customer? Sixth, residual risk: are gaps and partial controls described plainly with remediation actions? Seventh, operational assurance: does the vendor show how controls are monitored after the report is written?

This is where IRAP-style thinking beats the old compliance theater. A vendor can have a certificate and still leave the customer exposed if the relevant service was not in scope, the integration pattern changed, the AI feature was added after the assessment, the evidence is stale, the region is different, or the customer's own configuration carries most of the risk. A scoped assessment forces those uncomfortable details into daylight.

The Australian cloud guidance also makes a point that every global security team should borrow: cloud consumers remain accountable for their own data. Even if a cloud service provider has been independently assessed, the customer still has to configure identity, logging, monitoring, encryption, network controls, data governance, and incident response correctly. In AI systems, the same shared responsibility shows up as prompt governance, tool permissions, model gateway policy, audit logging, evaluation data protection, and human approval paths.

For CISOs, the best use of an IRAP-style score is not procurement theatre. It is decision support. Ask the security team to turn every vendor assessment into a short heat map: green for evidence-backed controls that match the intended use, amber for controls that need customer configuration or compensating control, red for gaps that block launch, and grey for areas not assessed. The grey items are often the most important because they reveal what the badge did not cover.

For AI vendors, this is also a marketing opportunity, but only if done honestly. Instead of saying 'we are secure', publish a trust package that explains assessment scope, AI feature boundaries, model and data-flow diagrams, customer responsibility, supported logs, retention settings, admin controls, incident process, and known limitations. Buyers trust clear boundaries more than heroic claims.

For security teams reviewing AI coding assistants, model gateways, agent platforms, MCP servers, or autonomous SOC tools, the IRAP-style approach should become the default. Score the product by what it can prove. Can every agent have a distinct identity? Can every tool call be logged? Can write actions require approval? Can secrets be masked? Can egress be restricted? Can customer data be excluded from training? Can administrators export evidence? Can the product survive a prompt-injection exercise?

The future of security assessment will probably look less like one annual questionnaire and more like living assurance. Reports will still matter, but teams will increasingly ask for machine-readable control evidence, continuous monitoring, audit APIs, configuration drift signals, and event-level proof. A static report says what was true at assessment time. A useful assurance program shows what changed after Tuesday.

The HackWednesday take is simple: IRAP-style assessment makes sense because it respects reality. It does not reduce security to a sticker. It pushes teams to ask about scope, evidence, boundaries, controls, residual risk, and accountability. In the age of AI agents and cloud control planes, that is the conversation security leaders actually need.

Source notes

Every Wednesday post should link back to primary reporting or documentation so readers can verify claims quickly.