Free AAIA AI Operations Practice Questions
This 10-question domain test represents 46% of the ISACA AAIA exam and covers the AI lifecycle, data and model management, deployment, monitoring, change management, and operational controls. Work through each question, then review the explanation to identify the audit principle or control behind the answer.
Return to the AAIA practice-test hub or try the 10-question mixed test.
10 Sample Questions with Answers
Sample Question 1 — AI Operations
A retail bank deployed an updated AI credit-risk scoring model before a seasonal lending campaign. The change ticket exists, but independent validation sign-off was dated after production deployment. Management states the release was urgent and that early default-rate indicators improved. What should the auditor evaluate FIRST?
- A. Whether the deployment was linked to documented approval and exception handling before release (Correct answer)
- B. Whether post-release portfolio metrics indicate improved model performance after deployment
- C. Whether the model repository contains release notes for the deployed model version
- D. Whether the rollback procedure was referenced in the release planning materials
Correct answer: A
Explanation: A is best because the first audit priority is to confirm whether the model was authorized for production or released under a documented exception before deployment. Without that evidence, accountability and control operation are unproven. B is wrong because improved outcomes do not substitute for pre-release authorization. C is wrong because release notes support traceability, not approval. D is wrong because rollback readiness is relevant, but it is secondary to whether the release was properly authorized.
Sample Question 2 — AI Operations
An insurer uses an AI propensity model to prioritize customer retention offers. After an upstream feed change, the nightly scheduler logs show successful job completion, but the number of customers receiving scores declined. Management notes that campaign conversion rates for scored customers remain acceptable. What is the PRIMARY audit concern?
- A. The absence of source-to-score reconciliation may allow silent exclusion of records (Correct answer)
- B. The stable conversion rate may not reflect the quality of model predictions
- C. The updated lineage document may not describe all upstream data dependencies
- D. The scheduler completion logs may not identify the responsible support team
Correct answer: A
Explanation: A is best because the key operational risk is silent omission of source records after the upstream change, and source-to-score reconciliation is the control that would detect that completeness failure. B is wrong because acceptable conversion rates for scored customers can mask missing customers who were never scored. C is wrong because lineage documentation explains dependencies but does not detect excluded records in daily operation. D is wrong because support-team identification affects accountability, not whether the scored population is complete.
Sample Question 3 — AI Operations
A health insurer uses AI to triage claims for expedited or manual review. Monthly governance minutes show aggregate precision and recall within tolerance. An analyst's ad hoc review, however, found increased false negatives for one provider type after a policy change. No approved threshold exists for subgroup variance. Which control deficiency is MOST significant?
- A. Monitoring relies on aggregate metrics without required segmented thresholds (Correct answer)
- B. Governance minutes do not include the analyst's ad hoc review file
- C. Operations has not quantified the cost impact of the subgroup variance
- D. The monthly review cadence may be too infrequent for claim triage
Correct answer: A
Explanation: A is best because a monitoring control that only reviews aggregate metrics cannot reliably detect subgroup deterioration, so the control is ineffective by design. B is wrong because missing documentation is secondary to the larger failure to require segmented monitoring. C is wrong because cost analysis may help prioritize remediation, but it does not address detection of the risk. D is wrong because more frequent review would still miss the issue if the dashboard omits the necessary subgroup measures.
Sample Question 4 — AI Operations
A fintech operates a real-time fraud scoring platform. Due to a small team, senior data scientists can develop models, approve production releases, deploy changes, and modify monitoring alert thresholds. Quarterly access reviews are performed, and privileged activity is logged. Which control deficiency is MOST significant?
- A. Incompatible privileges allow end-to-end model changes without independent control (Correct answer)
- B. Quarterly access reviews may not detect excessive privileges quickly enough
- C. Privileged activity logs may not be reviewed by the business owner
- D. Monitoring threshold changes may not be reported in release summaries
Correct answer: A
Explanation: A is best because the structural segregation-of-duties conflict allows one person to control development, approval, deployment, and monitoring changes, undermining multiple downstream controls. B is wrong because review frequency is a secondary issue when the role design itself is incompatible. C is wrong because logging is detective and does not prevent misuse of excessive access. D is wrong because release-summary reporting is narrower than the broader end-to-end privilege conflict.
Sample Question 5 — AI Operations
A consumer bank uses a third-party LLM API to draft customer service responses. The provider supplies a control report and consistently meets uptime commitments. The bank retains limited request-response logs and has not tested manual fallback during provider outages. Which risk is MOST relevant?
- A. The bank may lack traceability and continuity despite vendor service assurances (Correct answer)
- B. The provider may change its underlying model without improving response quality
- C. The service desk may underreport low-severity errors in customer responses
- D. The agents may rely too heavily on drafted language during peak periods
Correct answer: A
Explanation: A is best because the stated control gaps are internal: the bank may be unable to reconstruct AI-assisted interactions or continue operations if the provider fails, even if vendor SLA metrics are strong. B is wrong because model-change quality risk is plausible but not the main deficiency described. C is wrong because error underreporting is less directly supported than the missing traceability and fallback evidence. D is wrong because agent reliance is relevant operationally, but it is not the primary third-party dependency risk in this scenario.
Sample Question 6 — AI Operations
A payments company retrains its fraud detection model monthly. The pipeline automatically deploys the candidate with the highest fraud-capture uplift. Management cites improved results, but independent validation records are unavailable for recent versions. Which audit procedure is MOST appropriate?
- A. Trace recent promoted versions to validation evidence and approval before deployment (Correct answer)
- B. Compare monthly fraud-capture trends before and after each retraining cycle
- C. Review pipeline execution logs to confirm retraining completed on schedule
- D. Inspect model inventory records for version names and deployment timestamps
Correct answer: A
Explanation: A is best because the central audit issue is whether each retrained model reached production only after independent validation and formal approval. B is wrong because improved or changing fraud-capture trends are outcome evidence, not proof of controlled promotion. C is wrong because successful pipeline execution does not establish authorization. D is wrong because version inventory supports traceability, but it does not demonstrate that approval and validation controls operated before deployment.
Sample Question 7 — AI Operations
A healthcare revenue cycle team uses AI to suggest clinical coding before billing. Policy requires coder review of every suggestion, and workflow logs show the review screen cannot be bypassed. However, 98 percent of suggestions are accepted unchanged. Which evidence BEST supports the conclusion that human oversight is operating effectively?
- A. Quality review results showing sampled acceptances were assessed against coding guidance (Correct answer)
- B. Workflow configuration records showing coders must open each recommendation
- C. Management statements that experienced coders would identify obvious errors
- D. Monthly dashboard metrics showing the unchanged acceptance rate by coder team
Correct answer: A
Explanation: A is best because quality review of sampled acceptances against coding criteria is direct evidence that the human review step was substantive and aligned to standards. B is wrong because it shows the review screen exists, not that reviewers exercised judgment. C is wrong because management attestation is weaker than tested operating evidence. D is wrong because acceptance-rate metrics may signal behavior, but they do not prove the accepted recommendations were reviewed appropriately.
Sample Question 8 — AI Operations
A large retailer replaced an older demand forecasting model before a peak sales period. The previous model was retired after cutover. When forecast volatility increased, operations stated they could revert if needed, but no recent rollback exercise was available. Which evidence BEST supports the conclusion that rollback readiness is effective?
- A. Test results showing the prior approved model was restored from retained artifacts (Correct answer)
- B. Release templates showing rollback steps are included for production deployments
- C. Management confirmation that source files remain available for reconstruction
- D. Monitoring reports showing volatility has not exceeded escalation thresholds
Correct answer: A
Explanation: A is best because successful restoration of the prior approved model from retained artifacts is the strongest evidence that rollback can be executed in practice. B is wrong because documented rollback steps show intent, not tested capability. C is wrong because management belief and source-file availability do not prove recoverability within required timeframes. D is wrong because current monitoring status does not demonstrate that rollback would work if triggered.
Sample Question 9 — AI Operations
An insurer uses a third-party AI service to triage claims. The vendor provides monthly service availability reports, while internal teams retain responsibility for final claim decisions. During audit fieldwork, the auditor finds that model output quality issues are discussed in vendor meetings but are not tracked in the insurer's incident management system. What is the PRIMARY audit concern?
- A. AI incidents may not be escalated through accountable operational response processes (Correct answer)
- B. Vendor availability reporting may not reflect the insurer's business continuity needs
- C. Human claim reviewers may rely excessively on third-party model recommendations
- D. Service-level objectives may not include complete measures of production performance
Correct answer: A
Explanation: A is best because the evidence shows known output quality issues are being handled informally outside the insurer's formal incident process. That creates the clearest control failure in escalation, accountability, tracking, and remediation. B is a separate third-party resilience concern, but the finding is about quality issue handling rather than availability. C is a plausible human oversight risk, but it is not the issue demonstrated by the evidence. D is a possible control design weakness, yet the more immediate and supported concern is that actual AI quality incidents are not being routed through formal incident management.
Sample Question 10 — AI Operations
A global manufacturer relies on an AI-based predictive maintenance system. Policy requires model rollback when production monitoring shows severe performance degradation. Management states that rollback procedures were tested during the year, but no severe degradation event occurred in production. Which audit procedure is MOST appropriate?
- A. Inspect rollback test records and verify results against defined recovery criteria (Correct answer)
- B. Review production monitoring reports to confirm no severe degradation events occurred
- C. Interview maintenance managers about whether rollback procedures are understood
- D. Compare current model performance metrics with predeployment validation results
Correct answer: A
Explanation: A is best because, when no production trigger event occurred, the strongest evidence of rollback readiness is documented testing evaluated against predefined recovery criteria. That directly addresses operating effectiveness of the rollback control. B only confirms that the trigger condition did not occur; it does not show rollback capability. C provides indirect evidence of awareness, not proof the procedure works. D may help assess degradation risk, but it does not test whether rollback procedures would function as required.
AAIA Practice Test FAQs
What is the ISACA Advanced in AI Audit (AAIA) exam?
AAIA is an ISACA certification exam focused on auditing artificial intelligence systems. It covers AI governance and risk, AI operations, and AI auditing tools and techniques.
How many questions and how much time does the AAIA exam have?
The AAIA exam configuration is 90 questions in 150 minutes across three domains.
What score is required for AAIA?
ISACA uses a scaled score, with 450 on an 800-point scale listed as the passing score in the exam configuration. Practice-test percentages are not equivalent to an official ISACA scaled score.
How should I use AAIA practice questions?
Use mixed questions to find broad knowledge gaps, then use domain practice to review the concepts and explanations behind incorrect answers. FlashGenius readiness thresholds are study guidance, not official ISACA passing scores.
Which AAIA domain has the greatest weight?
AI Operations has the largest listed weight at 46%, followed by AI Governance and Risk at 33% and AI Auditing Tools and Techniques at 21%.
Explore AAIA Tests