Skip to content

Repository files navigation

badge-labs

Innovate.DTCC: Industry-Powered AI Hackathon Supported by FINOS

INSIGHT

Project Details

An AI-powered web application that transforms transit agency oversight reports into structured, actionable audit intelligence.

The Problem

Transit audit teams across the country — at agencies like WMATA, MTA, DOT, and Sound Transit — produce dozens of oversight reports each year. These PDF reports contain critical findings about financial risks, procurement weaknesses, safety gaps, and fraud exposure. But extracting usable intelligence from these documents is a manual, time-consuming process. Auditors must read each report, identify findings, classify risks, map them to fraud frameworks, and design internal controls — all before they can begin planning their next audit cycle.

There is no centralized, structured way to compare findings across agencies, identify recurring themes, or translate audit observations into concrete control recommendations.

The Solution

The INSIGHT = Industry Network for Systemic Issues, Gaps & Historical Trends uses AI to read audit reports and do the heavy lifting that normally takes auditors weeks of manual work. Here's what happens when a report enters the system:

  1. Upload a PDF audit report — The system accepts reports from any of the four supported transit oversight organizations (MTA OIG, WMATA OIG, DOT OIG, Sound Transit).

  2. AI reads and extracts findings — Using AWS Bedrock (Claude Sonnet 4), the system reads the full document and identifies every audit finding, including the risk it exposes, the root cause (based on what the source document actually says, not inferred), the auditor's recommendation, and a direct evidence quote from the report.

  3. Findings are classified — Each finding is assigned to one of 9 standardized audit categories (such as Procurement / Contracting, Finance / Grants / Funds Management, or Cybersecurity / Information Security), making it possible to compare findings across agencies and years.

  4. Fraud risk is assessed — The system evaluates whether each finding has fraud exposure and, if so, maps it to the ACFE Fraud Tree — the industry-standard framework from the Association of Certified Fraud Examiners. This gives auditors a three-level fraud classification (e.g., Corruption > Conflicts of Interest > Purchasing Schemes) with a confidence rating and rationale.

  5. Internal controls are suggested — For every finding (not just fraud-related ones), the system generates 1-3 suggested controls written in structured audit language:

    • WHO performs the control (specific department or role)
    • WHAT/HOW the control works (written in active voice with clear frequency and method)
    • EVIDENCE that proves the control is operating (specific documents or records)
    • Each control is labeled with one of 5 control types:
      • Preventive — stops the risk event from occurring
      • Deterrent — discourages the risk event
      • Detective — identifies the risk event after it occurs
      • Corrective — remediates the impact after detection
      • Compensating — mitigates risk when primary controls are infeasible
  6. Interactive dashboards support planning — Visual charts show risk distribution across categories, control gaps, trend analysis by agency, and other metrics that help audit teams prioritize their annual audit plan.

What Makes This Different

  • Source-grounded analysis — Root causes and evidence quotes come directly from the audit document, not from AI speculation.
  • Controls for every finding — Not just fraud scenarios. Every finding gets actionable control suggestions with clear ownership and evidence requirements, each labeled with one of 5 control types (Preventive, Deterrent, Detective, Corrective, Compensating).
  • Cross-agency visibility — Compare findings and trends across MTA, WMATA, DOT, and Sound Transit in one place.
  • Audit-ready language — Controls use the WHO/WHAT-HOW/EVIDENCE structure that auditors actually use in their work.
  • ACFE Fraud Tree integration — Fraud risk alignment uses the globally recognized framework, not a custom taxonomy.

Application Pages

Page What it does
FactPack The main dashboard. Three tabs: Findings (expandable cards with risk, root cause, recommendation, evidence quotes, and suggested controls with 5 control type badges), Fraud Scenarios (ACFE-structured table with filterable columns, expandable rows, perpetrator types, and confidence levels), and Suggested Controls (WHO/WHAT-HOW/EVIDENCE format with control type labels).
Planning Insights Four dashboards: Top Audit Areas by Finding (pie chart with drill-down), Emerging Risk Trajectories (recent vs. prior half comparison), Top Risk in Focus vs IIA 2026 Global Risk Report (benchmark comparison), and Control Recommendations by Risk Category (stacked bar with all 5 control types).
Data Hub Manage data sources, upload new PDF reports for AI analysis, toggle agencies on/off, and track sync status.
Ask Engine AI-powered query interface backed by AWS Bedrock. Supports natural language questions about the audit database and PDF file uploads for automated risk register and control matrix generation.

Tier 1 Audit Category Taxonomy

Findings are classified into 9 standardized categories:

Category Covers
Human Capital / Workforce Staffing, training, labor, HR, overtime, succession planning
Information Technology (IT) Systems, applications, data management (non-security)
Cybersecurity / Information Security Network security, access controls, incident response, cyber policies
Procurement / Contracting Bids, contract management, vendor oversight, change orders
Finance / Grants / Funds Management Grants, financial controls, billing, cost allocation, FTA compliance
Health, Safety & Security Operational safety, physical security, safety management systems
Operations & Maintenance Rail/bus operations, maintenance, reliability, asset management
Governance / Risk / Compliance Enterprise risk management, policies, oversight structure, ethics
Other Findings that don't fit the above categories

Tech Stack

Layer Technology
Frontend React, TypeScript, Vite, Tailwind CSS v4, Shadcn UI, Recharts
Backend Express.js, TypeScript
Database PostgreSQL with Drizzle ORM
AI/ML AWS Bedrock (Claude Sonnet 4) for PDF analysis
State Management Zustand (filters), TanStack Query (server data)
Routing Wouter

How the AI Pipeline Works

PDF Upload
    |
    v
Text Extraction (pdf-parse)
    |
    v
AI Analysis (AWS Bedrock / Claude Sonnet 4)
    |
    v
For each finding, the AI extracts:
  - Title, Risk, Root Cause, Recommendation
  - Tier 1 Category classification
  - Evidence quote from source document
  - ACFE Fraud Tree scenario (when applicable)
  - Suggested internal controls (WHO / WHAT-HOW / EVIDENCE + Control Type: Preventive/Deterrent/Detective/Corrective/Compensating)
    |
    v
Stored in PostgreSQL --> Served via REST API --> Rendered in React dashboard

Project Structure

├── client/                  # React frontend
│   ├── src/
│   │   ├── pages/           # FactPack, PlanningInsights, DataHub, AskEngine
│   │   ├── components/      # FilterPanel, Sidebar, shared UI
│   │   └── lib/             # API client, Zustand store, types
│   └── index.html
├── server/                  # Express backend
│   ├── routes.ts            # REST API endpoints
│   ├── storage.ts           # Database operations (Drizzle ORM)
│   ├── pdfExtractor.ts      # AWS Bedrock PDF analysis pipeline
│   ├── seed.ts              # Initial audit data seeder
│   └── db.ts                # Database connection
├── shared/
│   └── schema.ts            # Drizzle schema + Zod validation
├── demo/                    # Demo materials and screenshots
└── README.md

Running the Application

npm install
npm run dev

The application starts on port 5000 with both the API server and Vite dev server.

Environment Variables

Variable Description
DATABASE_URL PostgreSQL connection string
AWS_ACCESS_KEY_ID AWS credentials for Bedrock
AWS_SECRET_ACCESS_KEY AWS credentials for Bedrock
AWS_REGION AWS region with Bedrock access
AWS_SESSION_TOKEN AWS session token (for temporary credentials)

Team Information

ACT!vated Intelligence is composed of representatives from the WMATA Office of Audit & Compliance, operating under the Audit & Compliance Technology Center of Excellence (ACT!). The team name reflects the group's mission: activating intelligent technology to modernize how audit professionals plan, analyze, and act on oversight findings.

Team Members:

  • Katrina Welch-Smith
  • Tal Ron
  • Mike Revelle
  • Tony Rasoamiaramanana

Using DCO to sign your commits

All commits must be signed with a DCO signature to avoid being flagged by the DCO Bot. This means that your commit log message must contain a line that looks like the following one, with your actual name and email address:

Signed-off-by: John Doe <john.doe@example.com>

Adding the -s flag to your git commit will add that line automatically. You can also add it manually as part of your commit log message or add it afterwards with git commit --amend -s.

See CONTRIBUTING.md for more information

Helpful DCO Resources


License

Copyright 2026 FINOS

Distributed under the Apache License, Version 2.0.

SPDX-License-Identifier: Apache-2.0

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages