AI RESEARCH · PRODUCT · INFRASTRUCTURE · COMMERCIAL DILIGENCE
Where Decision Accounting could add value to an AI stack—and what has not been tested.
Decision Accounting supplies a proposed pre-decisional record and review architecture. Steadfastly is a reference implementation. The public evidence supports inspectable software and a computational lesson about consequences; it does not establish production impact inside any named AI or cloud platform.
Canonical public record: https://decisionaccounting.org/applied-research/
Proposed architecture
Governed institutional memory
The proposed layer preserves contemporaneous evidence, authority, uncertainty, predictions, selected and rejected alternatives, affected stakeholders, system effects, and reconsideration conditions. Agent retrieval should expose committed records; agent writes should create reviewable drafts; humans or authorized processes commit, attest, and supersede. This can complement conversation history, logs, connectors, and compliance exports by preserving the decision object itself.
Candidate interfaces include ChatGPT Enterprise workspaces, connectors, MCP tools, workspace agents, and compliance exports. A useful integration would need to demonstrate tenant isolation, permissions, minimization, retention, human review, exact record identity, committed-record retrieval, draft-only agent writes, and supersession. The public site provides no OpenAI deployment, partnership, endorsement, product integration, or measured customer outcome.
Steadfastly M365 demonstrates one implementation of drafting, commitment, retrieval, review, and supersession in a Microsoft 365 environment. Live-tenant validation remains pending. Diligence should test Entra identity, Graph and Purview boundaries, retention, eDiscovery, audit export, least privilege, failure recovery, tenant separation, and operational cost before any production claim.
The core record boundary is provider-neutral: committed records are readable, agent writes remain drafts, and human review controls commitment and supersession. The public candidate reports no production Google Workspace, Gemini, Vertex AI, Claude, or Anthropic enterprise deployment and no measured integration benefit.
A research agenda for scheduling, grid constraints, cost, reliability, energy, and water.
A data-center or model-training decision can be represented as a versioned choice among schedules, sites, hardware, power contracts, cooling designs, reliability targets, and workload deferrals. A Decision Accounting record could preserve forecasts, rejected alternatives, authority, cost and reliability expectations, grid constraints, energy and water effects, affected communities, and reconsideration triggers. That is a testable application design. The public program currently has no field experiment, validated estimator, production scheduler, infrastructure benchmark, or admitted SAPM value establishing improvement in these domains.
Ready for a scoped test
Research question
Does a pre-decisional record improve reconstruction, constraint compliance, forecast calibration, cross-team review, or post-incident learning relative to existing tickets, logs, change records, and observability systems?
Existing stack first
Required comparator
Compare against the lab’s current scheduling, capacity-planning, incident, environmental, approval, and compliance systems. Measure incremental information, latency, operator burden, error detection, decision reversal, reliability, and resource outcomes.
Pre-specified
Rejection conditions
Reject or narrow the layer if it duplicates existing provenance, increases operational risk, creates unacceptable privacy or security exposure, cannot remain current, adds material friction, or fails to improve a defined outcome in a sufficiently powered evaluation.
PARTNERSHIP AND ACQUISITION DILIGENCE
What a research, product, infrastructure, legal, or commercial team should verify.
Reproduce the Calvano/DRCS experiment and inspect the null findings.
Threat-model record fabrication, omission, retrospective rewriting, privilege, privacy, retention, access, and cross-tenant leakage.
Map the proposed record to existing logs, provenance, governance, compliance, and incident systems to determine incremental value.
Test user burden and failure behavior in a live but non-production tenant before field deployment.
Verify ownership, licensing, contributor history, third-party dependencies, security posture, and customer or deployment claims.
Keep held numerical work and provisional theorem status outside valuation or product claims.