News & Insights

NIST AI 800-4 and the Missing Governance Layer

Post-deployment monitoring has entered the official vocabulary. The next question is what must become visible.

nistai-governancepost-deployment-monitoringinteraction-geometry

By Scott J. Gardner, Founder and Research Lead

In March 2026, the National Institute of Standards and Technology published NIST AI 800-4, Challenges to the Monitoring of Deployed AI Systems. Its central signal is straightforward: pre-deployment evaluation is valuable, but it cannot tell us everything we need to know about an AI system once that system is operating in the world.

Deployment is not the end of evaluation. It is the point at which controlled assumptions meet changing users, workflows, interfaces, institutions, infrastructure, incentives, and other AI systems.

NIST identifies three reasons that post-deployment monitoring is necessary:

  1. to validate that an AI system continues to operate reliably in real-world conditions;
  2. to track unforeseen outputs, distribution shifts, non-deterministic behavior, and drift; and
  3. to identify unexpected consequences of integrating AI into new or changing contexts.

That is an important federal recognition of an unresolved governance problem. It does not validate RightMinds or any particular monitoring method. NIST explicitly describes best practices, validated methodologies, and common terminology as nascent and fragmented. The report names the terrain. It does not claim that the instrument panel has already been built.

What NIST AI 800-4 establishes

NIST organizes post-deployment monitoring into six categories:

Monitoring categoryGoverning question
FunctionalityDoes the system continue to work as intended?
OperationalDoes it maintain reliable service across its infrastructure?
Human factorsIs the interaction transparent to people and are its outputs high quality?
SecurityIs the system resilient against attack and misuse?
ComplianceDoes it adhere to applicable laws, standards, controls, and directives?
Large-scale impactsWhat downstream effects emerge across people, institutions, and society?

The report also identifies cross-cutting obstacles: fragmented logging, difficulty detecting degradation and drift, limited research on human-AI feedback loops, underexplored methods for detecting deceptive behavior, weak incident-sharing systems, competitive pressure against oversight, monitoring burden, and the absence of trusted standards for methods and tools.

These are not merely data-collection problems. They are problems of what counts as the system, where its consequential behavior becomes observable, and who has enough visibility to contest the resulting claims.

A deterministic filter is not governance

A deterministic filter can regulate expression without governing behavior.

It can determine whether one output crossed a declared boundary. That may be necessary. It cannot, by itself, show how the interaction arrived there, what alternatives were progressively excluded, whether confidence rose faster than evidence justified, or which participant increasingly shaped the path.

Output checks ask: Did this response violate a rule?

Trajectory governance must also ask:

  • What changed across the interaction?
  • When did the change become persistent?
  • Which parts of the coupled system shaped it?
  • Did the range of viable options widen or narrow?
  • Can the evidence be reconstructed and challenged?
  • Is correction still possible?

The governed object is not only a model response. It is the coupled trajectory formed by models, people, memory, retrieval, tools, interfaces, policies, incentives, workflows, and time.

The missing interaction layer

Normal chat logs preserve words. They rarely make the shape of the interaction legible.

RightMinds uses interaction geometry as a public description of observable, trajectory-level patterns: how an interaction moves through meaning-space; whether it broadens or narrows; how strongly prior turns constrain later ones; and how influence shifts between participants over time.

The practical question at the center of this work is simple:

Who is steering?

That question is not answered by classifying a person’s beliefs or inferring a hidden psychological state. It can be approached through interaction-level evidence such as timing, turn structure, topic and frame movement, question-versus-assertion balance, options-versus-conclusions, stylistic convergence, and changes in which participant introduces or closes alternatives.

Candidate monitoring categories include drift accumulation, narrowing, steering balance, control transitions, uncertainty mismatch, and relational change. These are research questions, not claims of mind-reading, manipulation detection, psychological diagnosis, or moral certification.

A layered response

RightMinds approaches the monitoring gap at three different levels of maturity and scale.

RightMinds: the framework

RightMinds is developing a framework for interaction-level observability and governance. Its central claim is that consequential AI risk increasingly lives in trajectories, not only in isolated outputs. Publicly, the framework describes the categories that should become visible and the evidence needed to reconstruct change. Proprietary measurement methods remain outside that public description.

Orbital: the instrumentation primitive

RightMinds Orbital is a local-first, browser-based application for cross-platform AI conversation continuity and observability. It brings supported conversation histories into a user-controlled Thread Library, presents selected threads as a timestamp-ordered Turn View, and provides early interaction telemetry anchored to the turns it describes.

Orbital is designed around cognitive and data sovereignty. Conversation data remains local unless the user explicitly chooses to share selected context with an AI platform or with RightMinds support. Its product center is visibility: helping a person inspect how an interaction changed and see who appears to be steering without surrendering their history to another centralized service.

Orbital’s metrics are experimental. The application is an early instrumentation primitive, not a validated regulatory, diagnostic, or forensic instrument.

GDO: the observatory proposal

The Global Drift Observatory is a proposed independent observability layer for examining AI-mediated change across platforms, institutions, media, markets, and public systems. It extends the monitoring question from individual interactions toward NIST’s large-scale impacts category.

GDO remains an architecture and governance proposal. Questions of institutional independence, funding, access, contestability, correction, data governance, capture resistance, and public authority are prerequisites. An observatory that cannot be audited or contested would reproduce the opacity it claims to solve.

Crosswalk against the six NIST categories

NIST categoryRightMinds relationshipCurrent posture
FunctionalityInteraction telemetry can help reveal whether conversational behavior remains consistent with declared operation across time.Partial alignment. Not a general functionality test suite.
OperationalCross-platform continuity and provenance address parts of fragmented interaction logging.Outside primary scope. Infrastructure reliability remains a conventional observability domain.
Human factorsSteering balance, narrowing, drift, uncertainty mismatch, and relational change concern interaction quality and transparency.Strongest alignment; experimental. Validation remains incomplete.
SecurityTrajectory monitoring may expose behavioral changes or misuse patterns that merit investigation.Partial alignment. Not a cybersecurity replacement.
ComplianceReplayable, version-bound evidence could support review when mapped to a specific obligation.Future research. No certification claim is made.
Large-scale impactsGDO proposes independent observation of downstream AI-mediated change.Architecture proposal. Governance and methods are not mature.

The crosswalk also makes the scope clear. RightMinds contributes most directly to human-AI interaction monitoring. Established operational and security tools remain essential, while compliance and large-scale impact monitoring require additional methods, institutions, and validation. Trustworthy measurement begins with being explicit about what an instrument can and cannot show.

What RightMinds adds

RightMinds adds four emphases to the monitoring agenda.

Trajectory-level analysis

Some failures are cumulative and path-dependent. They become visible only when turns are examined as a sequence with direction, momentum, and changing constraints.

Cross-platform continuity

People increasingly reason across ChatGPT, Claude, Gemini, OpenWebUI, and other systems. Vendor-specific logs divide one person’s cognitive workflow into separate silos. User-controlled continuity is both a usability and an observability requirement.

Interaction-level provenance

A measurement claim should be traceable to the interaction segment, method version, configuration, and evidence that produced it. Friendly labels are not enough. Any externally relied-upon measurement must preserve the identity and limitations of the instrument that generated it.

Visibility without proprietary model access

Not every monitoring problem requires model weights, private chain-of-thought, or proprietary internal logs. Interaction-level signals can provide limited but useful evidence from the observable exchange itself. This does not make internals irrelevant; it creates an independent evidence surface when internal access is unavailable or institutionally conflicted.

The measurement boundary

Observability is not moral certification.

RightMinds can verify that the instrument worked. It cannot absolve the operator.

An intact measurement channel can show that evidence was captured and processed as declared. It cannot determine that the measured system is acceptable, lawful, ethical, or safe in every context. Those judgments require legitimate authority, domain expertise, affected-party participation, and the ability to contest both evidence and standards.

RightMinds also rejects using interaction observability for involuntary inference, covert surveillance, behavioral manipulation, or systems that reduce a person’s ability to exit and retain interpretive freedom. A tool designed to reveal steering cannot legitimately become a hidden steering system itself.

What remains unproven

The current RightMinds and Orbital metrics are experimental. They have not been peer-reviewed, standardized, certified for forensic use, or validated as evidence of manipulation, coercion, psychological state, or regulatory violation.

Required work includes stable and versioned method definitions, calibration, reproducible computation, benchmark corpora, confidence intervals, false-positive and false-negative characterization, sensitivity and specificity studies, cross-model replication, longitudinal testing, adversarial robustness, and independent peer review.

The institutional work is equally important. Independent monitoring requires credible governance, funding, access rules, privacy protections, correction procedures, and mechanisms through which affected people can challenge findings.

NIST AI 800-4 does not endorse RightMinds, Orbital, GDO, or their methods. The alignment is between NIST’s documented monitoring needs and the problem space RightMinds is investigating.

Visibility is the first requirement for governance

Governance requires more than the ability to block an output. It requires legitimate, observable, accountable, contestable, and corrigible shaping of consequential system behavior across time.

We cannot govern what only the platform can see. We cannot contest a measurement whose method and provenance disappear. We cannot correct a trajectory if we notice it only after its assumptions have hardened into infrastructure.

NIST has named the monitoring vacuum. RightMinds is building early instrumentation for one part of it: the interaction layer where human and system behavior become coupled, cumulative, and consequential.

The next task is not to pretend the instrument panel is finished. It is to build it carefully enough that the measurements remain visible, bounded, reproducible, and open to challenge.