Siarhei Mardovich

Incline Trust

An AI exception-review workspace for corporate spend. It keeps model mechanism, uncertainty, and peer context visible, so a controller can accept, reject, tune, or escalate each flag in under a minute, then sign a record that survives review.

  • Concept prototype
  • Finance operations
  • AI trust
  • Exception review
  • Reg-tech
  • Model risk
  • SR 11-7
  • D3.js
  • Material Design 3

The 2 failure modes

01 · Problem

Controllers rubber-stamp a drifting AI score, or ignore it and fall behind at close. Either way the audit signature is theirs. The problem is that the model’s confidence, drift, and peer context are invisible at the moment of attestation, so the reviewer cannot defend the call they just signed.

These are the failure modes the workspace is designed against: hidden confidence, invisible drift, weak peer context, and a signature that arrives after the reviewer has lost the evidence.

AI finance operations has a specific failure pattern. The model narrows a high-volume queue, but the reviewer still carries the audit signature. Too much trust lets drift pass under a clean-looking score. Too little trust turns every case back into manual review.

So the real failure is the interface around the model. A recommendation without visible mechanism asks the reviewer to choose between obedience and suspicion. Neither is a reviewer.

The 5 questions the workspace answers

02 · One signing path

The screen is a single workspace that carries a case from filtered inbox to attested artifact in one session. Each visualization answers a distinct question the reviewer must answer before signing the record.

6 Coordinated views: queue, score, timeline, weights, peers, and audit, in one workspace.
<60s Inbox to signed record on the designed path with no context switch. A design target, not a measured time.
4 Tunable signals in the weight panel, rather than 12.
1 Immutable record per decision: signals, weights, sigma band, peer cohort, timestamp, reference ID.

The guided path moves from queue to explanation to weights to score to timeline to verdict. That order matters. Reading the explanation before touching the weights stops the reviewer tuning toward a preferred outcome. Reading history before signing stops a decision that ignores seasonality or peer movement. The questions below follow that path.

  1. Weights. Is the score driven by signals the reviewer trusts, or by a signal that should be argued with, and how sensitive is the output to each one?

  2. Score. Is the point estimate stable, or is the same score carrying enough variance to change scrutiny?

  3. Timeline. Does the current month break from its own history, or only look strange in isolation?

  4. Peers. Is this case the outlier, or is the cohort itself drifting into a policy question?

  5. Audit. What did the reviewer see, change, and attest at the moment the record was created?

The gauge keeps uncertainty visible · the timeline keeps history in a 12-month geometry · the sliders keep tuning explicit · every visual is accountable to the signature at the end of the path.

Fig. 01

Reviewer’s 60 seconds

5 frames across one session, from inbox to signed record: the confidence-band inbox with score and uncertainty visible before a case is opened, the decision-weights panel whose tunable sliders keep the mechanism legible, and the signed audit card every verdict emits.

The inbox ranks by confidence band, not point score

03 · Calibration

Score and uncertainty are both visible before a case is opened, so a wide band earns attention even when the point estimate looks calm. Ranking that way traded a familiar ranked list for a queue that surfaces variance, not just severity.

The queue is governed by thresholds instead of a flat ranked list, so policy can move without hiding the tradeoff. A separate threshold instrument makes that tradeoff tangible. Drag the escalation and auto-approve handles and the same synthetic queue re-bins in real time. The counts are a side effect. The purpose is to make policy changes visible before they become reviewer behavior.

The weight panel that sits behind the score exposes 4 tunable signals rather than 12. That traded model fidelity for a panel a controller can actually read inside a 60-second review.

Fig. 02 Exception review workspace

6 coordinated views on 1 case. Open the prototype and walk the signing path from inbox to signed record.

Fig. 03 Threshold instrument

Drag the escalation and auto-approve handles and the same synthetic queue re-bins in real time, so a policy change is visible before it becomes reviewer behaviour.

Data

Confidence scores in the threshold instrument are illustrative. The handles re-bin a synthetic queue.

Model-risk concerns mapped to screen controls

04 · Governance

The governance map translates SR 11-7 expectations into things a reviewer can see and use without exposing the proprietary model. Every expectation lands on a control that is already on the screen at the moment the question is live, rather than in a document the reviewer reads later.

Corporate spend is the setting. Any regulated queue where a person signs off on an AI flag is the same design problem: a model narrows the work, a person owns the call, and an audit reads the record later. AML transaction monitoring, KYC and KYB anomaly review, fraud detection queues, credit-risk review, and operational-risk exceptions all take that shape.

Scope

SR 11-7 governs institution-level model inventory and validation. This workspace covers the reviewer-level decision record only.

Every verdict path creates the same immutable record

05 · Accountability

The state machine is the accountability model behind the screen. A low-confidence case is flagged, routed to review, then approved, rejected, or escalated. The branch can change, but the end condition does not. Accept, decline, and escalate all terminate in the same immutable record, and escalated cases re-enter review first.

The audit card stays dark until the verdict is signed. That traded a more impressive first load for honesty, because a pre-signature audit is a design lie.

This is what a reviewer signs, and what the audit reads back later. Values below are illustrative.

Verdict
Accepted
Captured signals
recurrence_match, amount_deviation
Captured weights
0.70, 0.50, 0.80, 0.60
Sigma band
+/- 12
Peer cohort
Same category, outlier
Timestamp
2026-06-14T14:32:00Z
Reference ID
IT-C-04-2026-0614-001

The immutable record the audit review reads back.

How it was built, and what I would measure

06 · Craft
Stack and provenance
  • Concept prototype: single-file HTML on a Material Design 3 dark shell, D3.js v7 visualizations, no framework runtime.
  • D3.js views: confidence-band inbox, score gauge, 12-month timeline, weight-tuning controls, peer cohort fit bars.
  • Threshold calibration runs as a separate iframe instrument, so the D3 logic stays outside the page editor and can be updated independently.
  • Accessibility: native label-and-input pattern on the attestation checkbox, visible focus states, keyboard shortcuts for queue review, reduced-motion preference respected in both the prototype and the threshold instrument.
  • Frontend coded by directing AI coding agents, so the judgment is the design work and the keystrokes are accelerated.
  • Role: principal product designer and prototype engineer, drawing on 15 years designing decision interfaces across capital markets and enterprise data products. Designed the exception-review interaction model, the score-and-uncertainty system, the weight controls, and the attestation artifact.

Building it clarified that the hardest accountability surface is not the score but the record the reviewer signs, so the audit card drove the design, not the gauge.

What I would measure: median time from case open to signed verdict against the reviewer’s current queue tool; how often weights are adjusted before signing, which is the tuning-toward-a-preferred-outcome signal; and the share of signed records that survive a second-line audit challenge without rework. No baseline exists yet. The instrument is designed to capture all 3 from the first session.

Contact

Open to principal and staff product design roles. Greater New York City Area, remote or hybrid preferred, onsite flexible.