The false positive tax: alert fatigue and the math of analyst time
جھوٹے مثبت کا ٹیکس: الرٹ تھکاوٹ اور تجزیہ کار کے وقت کی ریاضی
35 min read
Three ways to see it
Precision and recall are the two numbers every fraud team must know cold. Precision is true positives divided by the total number of alerts. If 100 alerts are raised and only 5 turn out to be real fraud, precision is 5 percent. Recall is true positives divided by all real fraud, including the cases the system missed. If 100 frauds happened in a month and the system caught 60, recall is 60 percent. A team that maximises only recall produces an avalanche of alerts. A team that maximises only precision misses fraud. The honest tradeoff is captured by the F1 score, the harmonic mean, and by the precision-at-fixed-recall metric, which is more useful when your analyst capacity is the binding constraint.
Do the math for a typical Pakistani mid-tier bank. Suppose 10 million transactions per day. The current rule engine alerts on 0.5 percent, meaning 50,000 alerts per day. Triage at 90 seconds per alert needs 75,000 analyst minutes, or 1,250 analyst hours, or 156 analysts working an 8-hour shift just to look once at every alert. Most banks do not have 156 AML analysts and so most alerts are auto-cleared by simple rules without a human glance. The ML promise is concrete: hold recall constant and cut alert volume by 70 percent through better prioritisation. That moves you from 156 analysts to 47, or it lets your existing 47 analysts spend 4x longer on each alert. Either way, precision goes up, fatigue goes down, and missed fraud goes down with it.
Fatigue is not just about hours. It is a measurable cognitive degradation. Studies of radiologists, intelligence analysts, and yes, fraud analysts, find that detection accuracy on the same task drops 10-30 percent after the first ninety minutes of repetitive review and continues to drop through the day. The implication for fraud teams is direct: even if you have enough headcount, the alerts triaged in the last hour of a shift have systematically lower precision and recall than alerts triaged in the first hour. Two practical mitigations are documented to work. First, rotate analysts off triage every two hours into a different task such as case investigation or quality review. Second, sort alerts so that the highest-score, highest-stakes alerts arrive at the freshest analyst. AI gives you the score that lets you sort.
Quick check
Quick check: what makes modern AI different from a rule-based program?
The why-tree
Why-tree level one: why is alert fatigue not a willpower problem? Because the brain reuses the same pattern-recognition circuit for every alert, and that circuit physically tires through the day. No amount of training, motivation, or coffee can overcome the underlying neural cost. The fix has to be structural, not personal.
Try this with Claude
AI-edge prompt: 'My bank has 22 AML analysts handling 38,000 alerts per day with average precision of 4.2 percent. Each analyst costs PKR 180,000 fully loaded per month. A missed mule case costs PKR 280,000 on average. A false positive costs PKR 720 in analyst time and PKR 200 in customer service. Build me the unit economics model and tell me how much I can spend on a new ML monitoring layer that lifts precision to 18 percent at the same recall. Show your work and identify the assumptions most worth checking.' Compare the model's logic to your own quick estimate.
Sources
Sources and further reading. SBP Consumer Protection Framework 2024. SBP Banking Conduct Department circulars on customer impact reporting. Wickens, Hollands et al, Engineering Psychology and Human Performance, on cognitive vigilance decrement. Ruchi Sanghvi et al, on alert fatigue in security operations. Kahneman Thinking Fast and Slow on cognitive cost. ACAMS guide on tuning detection scenarios. Wolfsberg Group Statement on monitoring screening searching. The 2023 Federal Reserve study on alert fatigue in BSA AML compliance. Pakistan Banking Mohtasib annual report.