Skip to main content
From a 2% Sample to Every Filing Reviewed - Regulatory & Compliance case study
Public Sector · Regulatory & Compliance

From a 2% Sample to Every Filing Reviewed

AI-assisted review that reads every filing, appraises it consistently, and proves why, with a person deciding every case.

The Problem

Our client is a national public-sector regulator that maintains a register of tens of thousands of registered entities and receives more than 24,000 annual filings each year. Each filing is a structured return the entity has written about itself, a summary taken largely on trust, sitting on top of the real evidence: financial statements, assurance reports and performance information, filed as documents in hundreds of different layouts.

Behind the forms sits a small specialist review team. With finite capacity against that volume, only a small fraction of filings, historically in the order of 2%, could ever be reviewed in depth. The rest passed onto the register unread. Risk-based targeting helped, but the underlying problem remained: a team can get across a sample, but risk does not confine itself to the sample. The evidence base for regulating an entire sector was, in effect, a rounding error of the sector itself.

The consequence was a quiet accumulation of risk: incomplete or inconsistent filings entering the register undetected, non-compliance surfacing only once it became a public problem, and scarce specialist expertise consumed by mechanical checking rather than judgement.

The Mahi

DataSing partnered with the regulator to deploy Sentinel, an AI-powered compliance-review platform built on Azure OpenAI in the New Zealand region, behind government single sign-on. Sentinel is built on one principle: read everything, appraise it consistently, and prove why, while a person decides every case. It pairs two engines over a single dataset.

The individual-case engine reads every filing. On arrival, an acceptance gate confirms the submission is the right document (complete, signed, not a draft, and lodged by the declared entity), catching broken filings in plain language on day one rather than months later. Each accepted filing is then assessed against a statutory rubric of obligations and criteria: every criterion receives a rating, a quoted passage of evidence with its page reference, the reasoning behind it, and a confidence level. Critically, the figures an entity declares on its return are reconciled against the figures Sentinel extracts from the filed accounts, and where the form and the accounts disagree, the divergence itself becomes a signal. Each filing is then triaged (auto-pass, review or escalate) under thresholds and weights the regulator sets and can publish.

The rubric is configuration, not code. Obligations, criteria, rating scales, weights and thresholds are defined per regime and versioned, so calibration is a policy decision the regulator owns rather than a change request to a software vendor.

For the analyst, expert time moves from finding problems to judging them. The day starts with a queue ordered by what matters most, escalations first and auto-passes collapsed. Each case presents the assessment obligation-by-obligation with the evidence alongside it, so “show me where” is a button rather than a request to a colleague. The analyst makes the decision, every action lands in an immutable audit trail, and feedback letters draft themselves from the findings.

The second engine turns every assessed filing into intelligence about the whole population, the first time the regulator can see the entire system on one screen. It answers questions no sample ever could: where the population actually struggles to comply and why, which obligations fail most in which parts of the sector, whether a guidance change actually moved the numbers, which entities are running out of financial road, and, for any single entity, its complete provenance-labelled history in one click.

Several design choices do the quiet work. Sentinel reads every page of every filing rather than skimming, so page 87 gets the same attention as page 1. It judges quality rather than mere presence, assessing whether a disclosure is meaningful against a rubric the regulator controls. It forgives the form, reading hundreds of different layouts equally well. And because every document read becomes structured, queryable data, it feeds the regulator’s existing analytics the ninety percent of the record they could never see before.

Responsible AI is engineered in, not bolted on. A named person makes every decision; nothing is enforced by the platform and there are no automated actions anywhere in it. Every figure on screen carries a badge declaring its source, and every AI judgement carries a confidence level, so “how did it know that?” always has an answer. Risk weights, thresholds and triage rules are set by leadership, in the open, and published; the algorithm is policy, not a quiet code change. Some things the platform could do, it deliberately does not: capabilities with privacy implications wait for a proper decision by leadership, and until then simply are not built. All data stays sovereign, in the New Zealand region behind government single sign-on.

The solution was delivered iteratively: proving accuracy against expert judgement on a controlled sample first, then running alongside the live process, before extending toward review at the point of filing.

The Outcome

In one assessment from the demonstrator dataset, an entity’s annual return declared total revenue of $773,475 while the filed statements showed $454,850, a 70% divergence. Its declared accumulated funds of $1,942,164 sat against $103,950 in the accounts, a 95% gap. The name on the document did not match the register record, a second and independent signal on the same filing. Neither figure alone looks wrong; a dashboard built over the return sees a healthy organisation. The divergence is invisible until something reads both sides.

At population scale, around 14% of filings show at least one material variance of this kind, an entire class of signal that simply did not exist before. Each is escalated to a person with both figures, the page references and the history attached. The system says “inconsistency requiring review”. It never says fraud; until a person decides, it is a signal, not a finding.

For the first time, every filing is read (not a 2% sample, but the whole population, every cycle) and every judgement is evidenced and owned by a person. Consistency now comes from the rubric; authority stays with the analyst. Specialist time has shifted from mechanical checking to the judgement calls that need it, with the riskiest cases surfaced first and the evidence already attached.

Leadership can now see the whole system at once: where risk is concentrating, how it is moving, which entities are heading toward distress, and whether interventions are actually working, turning a register that was largely taken on trust into a living, evidenced picture of the sector. By combining AI reading and appraisal with expert oversight, the regulator has moved from sampling the past to seeing the whole system: faster, more consistently, and with every judgement explainable and human-owned.

Impact & Outcomes

100%
Of filings read, up from a 2% sample
24,000+
Annual filings assessed each cycle
~14%
Filings showing a material variance
Zero
Automated actions; a person decides every case

Services used in this engagement

Ready to discuss your project?

Let's explore how DataSing can support your organisation.