The challenge
Reflex Security runs AI-driven incident response tabletop exercises. After each exercise, AI agents produce findings about how the team handled the crisis, and those findings end up in front of the CISO, the board, and insurers.
Their Founding Product Lead role called for AI output that shows its confidence, provenance, limitations, and human review. Regulators and insurers want evidence that a response plan was tested, so a wrong AI finding in that evidence is a real business risk.
My role
I chose the problem, did the analysis, designed the workflow, and built the prototype myself, using public sources only: the company CEO's podcast interview (The Cyber Security Matters Podcast, Ep 74), the job description, and IBM's Cost of a Data Breach Report 2026.
How I analyzed the problem
I scored four candidate workflows on what the build had to prove, and on the risk of looking like a copy of the company's own product.
I had solved this problem before, in a different industry. On the Marine Corps MAKE program, I designed the explosive safety inspection system: inspectors spent 1 to 2 weeks on a base, then briefed the base captain with a final inspection report.
A CISO's after action report on an incident exercise is the same job. Both turn a stream of field observations into evidence-backed findings that one decision-maker reviews and signs. So I carried the core capabilities across: every finding tied to its evidence, a review step before anything is final, and a report generated on the fly from approved findings.
| Workflow | Hands-on design depth | Roles and permissions | AI trust | Crisis coordination | Copy risk | |
|---|---|---|---|---|---|---|
| A. Live decision and notification board | Strong | Strong | Weak | Strong | High | |
| B. After action review (chosen) | Strong | Strong | Strong | Medium | Low to medium | |
| C. Scenario setup intake | Medium | Medium | Medium | Weak | Medium | |
| D. Readiness evidence for board and insurer | Medium | Medium | Medium | Weak | Low |
The after action review was the only workflow where every AI trust property is the work itself, with three-plus roles and a draft, reviewed, published lifecycle. The finding categories come from the four recurring response gaps named on the podcast: decision authority, notification order, containment versus recovery, and tribal knowledge, plus communication breakdowns and capability gaps. I folded in one element of the live board: the decision timeline shows where decision authority shifted and where notifications missed their required window.
How I solved it
- 1
Make AI confidence auditable
Confidence is computed from a fixed rubric (how many events, how many people, whether plan text backs it up), not asserted. Every finding shows its evidence tagged Agent, Human, or Document, carries a written limitation, and keeps edits in its history.
- 2
Keep humans in control
Rejecting a finding requires a reason. A low-confidence finding cannot publish without a written override. Send for approval stays locked until every finding has a decision, and executives cannot see the draft until the CISO approves.
- 3
One record, four views
Facilitator, CISO, executive, and participant each see their own dashboard, KPIs, and slice of the decision timeline. Participants see only the findings that cite them and can add context.
- 4
Build with agentic AI pipelines
Claude agents generated each version as a single self-contained HTML prototype, built the matching Figma design system through the Figma MCP connector, and ran headless browser checks on every version before I reviewed it.
- 5
Iterate on every review
Each review became a ruling applied in the next version: computed confidence, role-based views, a timeline redesign, slate grey instead of red for low confidence, and a CISO step to resolve disputed findings. After interview feedback, I added remediation tracking, a readiness trend across exercises, and a product direction page.
Results
- Demoed the prototype at the group interview on October 1, 2026.
- Delivered a Figma file alongside the HTML: 37 color tokens with dark and light modes, 17 text styles, reusable components, and screens for every role.
- Extended the slice into a full loop after interview feedback: findings become owned remediation actions, readiness is tracked across exercises, and the next exercise targets repeated gaps.
Try the prototype
Fully clickable, with fictional data. Use View as at the top right to switch between the facilitator, CISO, executive, and participant views.
Tools and stack
- Claude
- Agentic AI pipelines
- Figma
- Figma MCP
- HTML / CSS / JavaScript
- Headless browser testing
- NIST SP 800-61
- NIST CSF 2.0
- MITRE ATT&CK
