Provable human oversight for agentic AI

Everyone says they have a human in the loop. We prove whether that human could actually have stepped in.

Your agents act in seconds, a busy human signs off afterwards, and the log looks clean either way. FluxAI shows when that sign-off could not have been genuine, so a clean log stops passing for real oversight.

Why now

EU AI Act

High-risk obligations: 2 December 2027

Some of the obligations have been deferred, others already apply today. Most organisations are preparing by writing policy. When the requirement lands, you will need to prove that a human could genuinely exercise oversight, and that proof has to be built in from the start. The deferral buys preparation time that runs out on a fixed date.

Who it is for

Built for teams putting agents into production, where a lapse in oversight carries legal or financial consequences.

An AI agent moves money, approves a case, or reaches into regulated data. FluxAI works in the environments where those actions are supervised.

The roles

CTOCISOplatform securityCROHead of MRMcompliancesecond line of defence

The sectors

BankingPensionInsuranceProfessional services

Runtime, not audit theatre

The rest of the marketFluxAI
The proof is itself an AI judgementThe proof is arithmetic an auditor can recheck
Checked at audit time, staticEnforced at runtime, continuous
Two reviewers, two conclusionsSame conditions, same answer, every time
See the two axes: maturity vs runtime

The architecture

DARMA: five layers, five enforcement mechanisms.

Every agent in production raises five questions the instant it acts. DARMA answers each one with an enforcement mechanism.

LayerThe question at the moment of actionEnforced by
DDelegationWho activated the agent?Human in Control
AAuthorizationWhat may it do?Human in Control
RRuntimeDid this action pass policy?Human in Control
MModel IntegrityHas the model drifted?Human in Control
AAccountabilityCan we prove what happened?Proof of Human Oversight

Direct Advisory

You work directly with the person who built the framework and the products.

Engagements are structured to match your operational reality, operating either through the FluxAI stack or as an independent external resource.

From 1,500 DKK/ hour

Scope and estimate are agreed before we start.

The products

Two ways to close the oversight gap.

Product · Enforcement

Human in Control

Oversight enforced while the agent acts.

AI agents act; the runtime decides whether to let them. Each action is checked against your policy before it happens, and a human approves where it matters. Authority never collapses onto whoever pressed run. The same control governs a payments pipeline, a case-management system, or a clinical workflow.

For CTO, CISO and platform security leads.

See the four guarantees+

Nothing crosses unchecked.

Every action an agent attempts is checked against your policy before it touches anything outside the boundary. The check runs on mathematical logic rather than model-based judgement.

The model never sees regulated data.

The model only ever works on data it is cleared to see. Regulated fields never reach it.

Every decision is recorded.

Every agent decision is captured and exportable in the format an auditor expects. When a regulator asks who authorised an action or what data it touched, the answer is a record, not a reconstruction.

When governance is unreachable, the agent stops.

Acting depends on governance being available. If it is not, the action is denied and logged. Failing closed is the default, with no silent path that keeps running unchecked.

Deployment is scoped to your environment and risk tier. Starts with a Honeypot Assessment.

Human in Control in practice+

Payments · transaction authorization

An agent at a financial firm proposes a 50M DKK transfer. The runtime checks your risk rules in real time, then either lets it move or routes it to a human approver. When a regulator asks who authorised the transfer, the answer is in the record, not in someone's memory.

Advisory · confidential client data

An agent drafts client correspondence at an advisory firm. The brief contains confidential content. The runtime keeps that content away from the model. When a regulator asks whether confidential data reached a third-party model, the answer is no, and the record proves it.

Product · Proof of Human Oversight

Proof of Human Oversight

Prove when oversight could not have happened.

A regulator cannot deterministically prove that a human made the right decision. But it can be proven when the human could not have reviewed the case at all. Proof of Human Oversight turns rubber-stamping into a fact you can catch. Each finding is tied to the regulation it satisfies, and compliance mapping for Article 14 and DORA is included as a deliverable rather than sold separately.

For CROs, Heads of MRM, compliance and 2nd line of defence in regulated environments.

For model risk

Model Risk Oversight

Proof of Human Oversight, delivered in the language a model risk function and a regulator use, across banking, pension and insurance. It catalogues your AI use-cases and assigns a risk tier to each decision based on how severe and how reversible the outcome is. From there it sets where a human must sign off, and maps the result to model risk management, the discipline regulated financial firms already govern their models by. It is the same evidence, shaped so a model risk function can use it directly.

Delivered within a Proof of Human Oversight engagement, scoped by the number of models and use-cases.

See how risk-tiering works, and test how your AI holds up when the regulator walks in
See what the proof covers+

Catch approvals that were never real.

When an approval could not have been a genuine review of the case, Proof of Human Oversight records it as oversight that did not take place. A determination, not a judgement call, so it holds up to an auditor.

Keep approval independent.

Sign-off sits with a competent, independent role, and the most sensitive decisions require more than one. An approval that is really self-approval does not hold.

Bind every approval to the regulation.

Every approval is recorded and mapped to the article it satisfies (Article 14, DORA, NIS2, GDPR). Evidence an auditor can cite, not a checkbox.

Catch the gaps before the regulator does.

Where the record does not hold together, the gap surfaces internally instead of in a regulatory case. You find it first.

How it starts

Gap analysis

Free scoping

A short call maps where oversight can be proven across your agents, and leads into a Honeypot Assessment as the first deliverable.

Diagnostic

Priced on scope

Full proof of human oversight assessment with an evidence record and Article 14 / DORA mapping as a deliverable. Scoped from the assessment.

Proof of Human Oversight in practice+

Two approvals that looked fine but that a regulator would not accept.

Bank · credit approval

An agent's credit recommendation is approved almost instantly. Proof of Human Oversight records it as oversight that could not have taken place, before it becomes a regulatory case.

Advisory · GDPR notification

The DPO is on holiday, so IT approves the GDPR notification to meet the 72-hour window. Proof of Human Oversight flags the sign-off and preserves who actually decided what.

How we work

Three steps from first conversation to production.

Three defined steps, and you can stop after any one and walk away with what you have.

Step 1 · Assessment

Honeypot Assessment

5 business days · 50,000 DKK

The audit tests your AI agents in a controlled environment to find where governance is missing and what would have to change. It sets the right scope and price for the pilot, and stands on its own as a baseline you can show your board or auditor.

Step 2 · Pilot

Pilot in your environment

90 days · in your environment

The chosen product runs in your environment for 90 days, in production, not a demo. Human in Control governs up to 5 agents; Proof of Human Oversight proves oversight for one legal entity. Either side can end at 90 days.

Step 3 · Production

Standard package or larger deployment

Full deployment · support optional

Full deployment in your environment. Standard scales Human in Control to 25 agents or Proof of Human Oversight to 5 legal entities. Larger scope is priced individually. Ongoing support and policy updates are agreed separately.

The founder

I have seen what happens when compliance fails from the inside.

My background is in Danish public administration, in financial control, master data and the rollout of AI into regulated operations. Across those systems one pattern held: the model is almost never what fails. The breakdown is in the oversight. Someone approves a decision they never had time to check, and the log cannot prove afterwards what happened. The ownership gap stays open until a regulator asks.

AI agents are entering production faster than organisations can build controls around them. I built FluxAI to close that gap at runtime, and I advise organisations directly, on top of the platform I built.

FluxAI is an acknowledged contributor to IMDA Singapore's Model AI Governance Framework for Agentic AI. I take part in the European standardisation work for AI.

Straight answers

What buyers ask.

Do you read our data or our AI's reasoning?+

No. We do not look into the content of your data, and we do not claim to read a model's hidden reasoning. We watch behaviour and boundaries instead. If an agent moves outside its mandate, or the trail breaks, we stop rather than guess.

What can you actually prove about the oversight?+

We cannot prove a human made the right decision. We can prove when the oversight could not have taken place at all. Research shows people follow incorrect AI suggestions between 80 and 90 percent of the time, so a busy human catching the error in the moment is no real safeguard. Our answer is to build oversight into the structure: the machine handles the time-critical decision and can always fall back to a safe state, while the human reviews the case afterwards in peace. That is what makes rubber-stamping measurable.

Does it replace our compliance team or our people?+

Quite the opposite. The point is to make your people's oversight real, by having the machine lay the evidence out while the decision itself stays with your own staff. In that way we move the human away from what it is bad at, namely catching errors under time pressure, and towards what it is good at, namely weighing a case in peace.

How does it fit our existing setup?+

FluxAI sits on top of what you already have, so you never have to rewrite your existing systems. It fits into the most common AI-agent setups.

Can you prove it in an audit?+

Yes, and that is in fact the whole reason the evidence trail exists. Every single decision and action is written down immutably, so that afterwards you can document exactly what happened and who approved it, right down to the timestamp. You move from a claim that you did the right thing to a trail that genuinely holds when an auditor goes through it.

Are you auditors or a vendor?+

A vendor. We build the infrastructure, and your own second line of defence or an independent auditor does the oversight. We never attest to our own work.

How do we start?+

With a bounded gap analysis of your setup, where we look at where your oversight stands legally today, and where the holes are. The analysis is concrete and quick, and you come away with something usable whether or not we end up going further together.

See how human oversight actually fails

See where you stand.

Together we walk through where your agents stand, which DARMA layer is missing enforcement, and whether Human in Control or Proof of Human Oversight is the right place to start.