Skip to content

AI Evidence Assessment

Every evidence file you upload or collect is read by the platform’s AI and judged against the SCF assessment objectives of the controls it supports. The AI’s answer is always a suggestion: a person confirms or corrects it, and only then does it carry full weight in the Evidence Quality score. This page explains what the AI assesses, when, what the result looks like, and what a reviewer does with it.

The platform assesses evidence at two levels.

LayerWhat it looks atWhere you see itRole
Window assessmentEvery file that belongs to an evidence item’s current collection window, assessed together as one portfolioEvidence detail page → Assessment objectives for this window and Your Window Review; Evidence workspace Dashboard tab → Awaiting confirmation queuePrimary. This is the verdict that feeds the review queue and the scores
Per-file assessmentOne file on its ownFile preview → AI Assessment panelDiagnostic. Use it to understand why a particular file did or did not help

A control that is evidenced by several items also gets a composite: the roll-up of its window verdicts. The Evidence Quality axis reads the composite where one exists, then the window verdict, then the per-file verdict.

When you upload a file by hand, the upload form asks for preparer assertions: facts about the evidence that only the person who produced it can know. They are recorded exactly as entered and the AI never writes or changes them.

FieldWhat to enter
Effective periodThe dates the evidence actually covers (a report for August is effective 1–31 August whatever day you upload it)
CoverageWhat the evidence covers: a system, a site, a team, all users
PopulationThe full set the evidence was drawn from, if it is a sample
Selection methodHow items were chosen: all, random, judgemental
Sampling basisWhy that sample is representative
IPE source systemThe system the information was produced by (information produced by the entity)
Extract run onWhen the report or extract was run
Report or filterThe report name, query or filter used
Completeness checkHow you know nothing is missing

Two of these change what the platform does. The effective period decides which collection window the file belongs to (see below) and is used as the collection date for Evidence Health. The rest are shown to the AI as context and to reviewers on the evidence page, and they travel with the file into audit engagements.

Each tracked evidence item has a collection frequency. The frequency sets a window: how far back the platform looks for files that count as the current collection.

FrequencyWindow
Real time2 days
Daily2 days
Weekly9 days
Biweekly16 days
Monthly35 days
Quarterly95 days
Semi-annual185 days
Annual370 days
On demand35 days

Each window is a little longer than the cadence so that a collection that runs a day or two late still counts. An evidence item with no frequency is assessed on a monthly window and the assessment says so in its findings.

A file belongs to the window when its asserted effective period overlaps the window. A file with no effective period asserted belongs when its upload date falls inside the window. Every window assessment records, for each file, which rule put it there, so a reviewer can see why a file uploaded in September is in August’s window (its preparer said it covers August) or why a file uploaded on time is not (its preparer said it covers an earlier period).

The AI is given the evidence item’s catalog description, the controls it supports and their assessment objectives (the SCF’s per-control list of what an assessor needs to see), the expected artifact types for those controls, how the evidence is collected and from which system, and the text of every file in the window. It answers with:

  • One finding per assessment objective, with an advisory designation and the IDs of the files it relied on. The designations are appears_satisfied, gap_identified, not_applicable and cannot_assess. They are deliberately not assessor terms: the AI is telling you what it appears to see, not passing an audit.
  • An overall status for the window: sufficient, partial, insufficient, unassessable (nothing readable reached the model) or insufficient_sample (fewer files arrived than the cadence expects, for example one run in a daily window).
  • Coverage findings: which artifact types are missing, which sources went quiet, and what the AI was not shown (see below).
  • Effective dates it can read from each file, kept separately from what the preparer asserted so the two can be compared.

A daily collector produces twenty or more files in a monthly window, and they are often identical. The platform sends identical content to the AI once and tells it which files it represents; findings cite the representative file. If a window is still too large, the platform counts every file for coverage but stops sending content once the text budget is reached, and adds a coverage finding listing the files that were counted but not read. A reviewer who thinks those files could change an answer should open them directly.

  • On arrival. A manual upload or a collector delivery queues a window assessment after a short debounce, so a batch of files arriving together is assessed once.
  • Nightly. A sweep at 04:00 UTC assesses any item whose window has moved on since its last assessment.
  • On request. Reassess Stale Windows on the Evidence Coverage by Window card of the Evidence workspace’s Dashboard tab re-queues every item whose window has gone stale, and the API and MCP tools can queue one item or every tracked item.
  • After a revision request. Requesting revision on a window (see below) queues a fresh assessment straight away; files that arrive later queue another through the on-arrival trigger.

A window that has already been assessed with the same set of files is not assessed again.

Open an evidence item. Its detail page shows the current window, the files in it and why each is there, the AI’s overall status, and Assessment objectives for this window: one row per objective with the AI’s designation, its reasoning and the files it cited.

You have two decisions to make, and they answer different questions.

Confirm or correct the AI’s designations

Section titled “Confirm or correct the AI’s designations”

This is the question “did the AI read the evidence correctly?”.

  • Confirm AI assessment records that you agree with every designation as it stands.
  • To disagree, change the designation on any objective, give a reason, and save. The panel marks the window Corrected rather than Confirmed, and your designations replace the AI’s in the scores.

Either decision is written to an immutable version row. Nothing overwrites it: a later decision adds a new version, and the panel’s history shows every version with who made it and when. The version history is also the audit trail an assessor sees.

This is the question “is this collection acceptable as evidence?”, and it lives in Your Window Review on the same page.

  • Approve accepts the window as the collection for that period.
  • Reject records that it is not acceptable, with review notes.
  • Request revision asks the preparer for more or better files. When they arrive, the window is assessed again and returns to the queue.

Where the organisation requires reviewer independence, a reviewer who uploaded every file in the window cannot approve it; someone else has to.

The Awaiting confirmation card on the Evidence workspace’s Dashboard tab lists every AI verdict that no person has confirmed or corrected yet, worst first, so gaps surface before satisfactions. The card is marked AI Advisory to make its status clear. It lists window verdicts, the primary layer, and each row opens the evidence item so you can decide there. Unconfirmed per-file verdicts are visible on each file’s preview and, for integrations, through the review-queue API’s file tier.

The Evidence Quality axis of a capability theme’s posture is built from these verdicts, and it takes the human decision seriously:

  • An unconfirmed AI verdict counts at half its status weight.
  • A confirmed or corrected verdict counts at full weight.
  • A composite counts as confirmed only when every window it folded in has been confirmed.
  • unassessable and insufficient_sample windows are reported separately and are never scored as failures.

Confirming verdicts is therefore the fastest way to move Evidence Quality: the same evidence, once a person has stood behind the AI’s reading, is worth twice as much to the score.

The per-file assessment is still there for diagnosis. Open any file’s preview and the AI Assessment panel shows that one file’s objective findings, with Confirm AI assessment and Correct designations working the same way as on the window, and its own version history. Use it when you want to know which file in a window is carrying an objective, or why a file was read as unassessable.

Once an evidence item has a window assessment, the older per-file Approve and Reject document buttons are no longer shown for its files; acceptance is decided on the window. Evidence that has not yet had a window assessment keeps them.

Windows are only as good as the frequency behind them. The Frequency Health tile on the main Dashboard compares each tracked item’s declared frequency with the cadence at which files actually arrive and counts the items that are misaligned, for example an item declared monthly whose collector delivers daily. Show details lists them with the observed cadence and lets you apply the suggested frequency, which resizes the window the next assessment uses. Low-confidence observations (too few uploads to judge) are counted separately.