A Valve Is Not a Read
We put Sentience Governor inside another team’s self-healing building, running on Flower. The governance model traveled without a fork. Operation meaning, profiles, records and approvals did not travel on their own.
The water agent proposed shutting a riser valve, and Sentience Governor recorded the action as a read.
Nothing about the agent was wrong. The leak was simulated, and the agent did exactly what it was built to do. The record was wrong because Governor had no way to know what h2o.shut_valve means. It classified the operation the way it classifies any tool it has not been told about: from the words in its name. “Shut” is not one of the words that mean write. The most consequential action in the building was typed as the least consequential kind of operation.
This is the engineering log of that encounter, and of the others that followed when governance built around software agents met an experiment in physical infrastructure on someone else’s runtime. The companion essay in Sentient Notes, The Shutoff Valve Conundrum, asks what this means for physical AI.
Prologue: someone else’s building, someone else’s runtime
Antibody was built in a day by a small team at the Flower Collaborative Agent Hackathon at Stanford. It is a simulated self-healing building: fourteen specialist agents, each watching one system (power, water, HVAC, fire, egress, the building control system and more), plus a coordinator. It runs as a single Flower AgentApp, locally or on Flower’s hosted SuperGrid. The hackathon code is public, and the app is on Flower Hub.
We integrated Sentience Governor as its governance layer. Sentience Governor evaluates agent actions against declared governance context and records governance evidence without blocking execution. We went in with a narrow question: what does governing agents that act on physical systems require, using Governor as it is today?
Entry 1: One record per agent
Every agent got its own Governor session, with an identity (antibody-h2o, antibody-coordinator), a declared objective and a declared scope. A specialist’s scope is its own system: the water agent declares ["h2o"]. Each session writes its own Sentience Agent Execution Record, so one scan produces fifteen records.
Actions follow a naming convention, <area>.<verb>_<object>. Governor treats the text before the first dot as the target system, so an action on another area’s system would be recorded as outside the agent’s declared scope. That check was exercised in tests; no live scan produced a cross-lane proposal.
Antibody’s autonomy model lived in its own code, not in Governor. Cleanups ran on their own. Self-corrections adjusted a system within ±15%, re-checked the result and escalated after two failed attempts. Problems they could not fix, or a drop in overall building health, were escalated to a person. Consequential actions, such as shutting a valve or tripping a breaker, were held until a person approved them. Governor implemented none of these levels. It recorded what each agent declared, what it did, and the flags its profile raised.
Entry 2: Integrating without a fork
Governor’s wrapper is shaped for MCP clients. Antibody’s agents were Python functions running in threads. A small adapter in the application bridged the two: every action passes through Governor’s wrapper and then runs. Governor itself was not forked or modified.
The adapter did need three of Governor’s private internals: its session lifecycle calls, because the public lifecycle assumes async with and the agents ran in threads; its per-session operation-inference hook (Entry 3); and the location its loader reads profiles from (Entry 4). They made the integration work, and they are the first product inputs this experiment produced.
Entry 3: A valve is not a read
Governor’s wrapper infers an operation’s type from keywords in the tool name: write, update and create mean WRITE; delete and remove mean DELETE; exec and run mean EXECUTE; anything else is a READ. For a building, that produced:
| Action | Recorded as | What it does |
|---|---|---|
h2o.shut_valve | READ | Stops water to part of the building |
hvac.reset_damper | READ | Moves a physical damper |
start_automation | READ | Schedules recurring future runs |
cyber.run_audit | EXECUTE | Reads the building control system’s configuration |
The error runs in the unsafe direction. Governor’s first-write check and its read-to-write boundary signal key off the operation type, so for the most consequential actions in the building neither could fire. The high-consequence pattern did fire, because our profile named the valve. But a record saying an action is both high-consequence and a read is inconsistent, and anything that trusts the operation type, now or in a future enforcement mode, trusts the wrong thing.
We did not fix this by teaching Governor about valves. Antibody gained a domain table: 28 rows, one per action, each declaring an operation type and an autonomy level. h2o.shut_valve is EXECUTE and needs approval; cyber.run_audit is READ. An action missing from the table is treated as a write that needs approval, so the default for an unknown physical action is caution, not a guess. The adapter swaps Governor’s inference hook for a lookup in that table.
The same declaration feeds both consumers:
building_domain.py ──► Antibody's approval gate (which level applies)
└──► each Governor session (which operation type to record)
Because the gate and the record read one declaration, they cannot disagree about what an action is. We learned why that matters within hours. Before the table, autonomy levels lived in the application and high-consequence patterns lived in the profiles, and one work-order action ended up in one list and not the other the same day. Two copies of “what is consequential” drifted before the hackathon was over.
Who should declare what an operation is? Not Governor, which cannot know what a valve is. Not the agent, because an agent that classifies its own actions is self-certifying, and application code asserting a type per call is the same problem one layer up. Our conclusion is that operation meaning belongs to the operator, declared in the profile as tool pattern to operation type, alongside the existing high-consequence patterns, and evaluated by Governor.
This is the second time Governor has met the problem. Version 0.3.2 added semantic classification of shell commands because name matching failed for Bash. Two independent instances justify declared operation types as a product requirement. One domain and 28 rows do not justify a general ontology of physical consequence.
Entry 4: Profiles in, evidence out
Locally, everything above worked: all fourteen agents reported within a minute, and the consequential actions were flagged live. Then we ran the same app on SuperGrid.
Every Governor summary from that hosted run reported no profile and zero flags. Governor behaved as designed: it found no resolution file, governed with no profile, did not fail, and truthfully recorded that none was bound. Nothing surfaced the missing profile. The run looked governed and was not; we caught it only by reading the summaries.
The cause was simple. Governor resolves profiles from ~/.sentience on the operator’s machine, and a hosted executor has no operator home directory. We bundled the profiles inside the Flower app and pointed Governor’s loader at them. A test now runs Governor with no home directory and checks the bundled profiles still govern. The next hosted runs flagged the consequential actions.
The other direction mattered as much. Records stay where the agent runs. On SuperGrid they were written on Flower’s host, out of the operator’s reach. Each closed session now reads back its own record and emits a summary (actions, operation types, notable flags, approval) as a Flower run event and as log lines, which the console reads. That is remote evidence export: selected summaries from one executor to one operator. The full record never left the executor, and federated evidence across many locations was not exercised.
Flower made this visible by making remote execution easy. Running one app locally and hosted exposed the two assumptions Governor had been making: profiles are on the machine where the agent runs, and records are read where they were written. Remote execution needs profiles in and evidence out.
Entry 5: Proposed is not executed
A Flower run is one message, so Antibody holds a consequential action in run-series state, and a person’s approval arrives as the next run in the series. One action is split across two runs, and Governor’s record had no way to represent the gap.
When Antibody holds the valve shut-off, the adapter records the proposal through the same call path as an execution, with a no-op function. The record shows h2o.shut_valve, EXECUTE, high-consequence, for a valve that was never shut. To tell a proposal from an approved execution, each approved action gets its own Governor session, whose declared intent carries an authorization claim, approved by <name> (<id>). Reading the two records together tells you proposed from approved. Neither says so on its own. The approval path was tested against real Governor sessions at the hackathon and carried through live afterward, in the public demo.
A third state is missing entirely: the result. Governor can record a proposed action and the authorization behind it, but that does not establish that the action was actually executed, or that it produced the intended effect.
The record needs a lifecycle: proposed, held, authorized by whom, executed, result.
Entry 6: Who authorized it
Governor records the authorization claim it is given; it cannot know whether the claim is true. In the hackathon version the approver’s name came from the browser, so anyone could type one, and the public demo now sets it from a verified session. An authorization claim is only as good as the path it arrived on, and the record should eventually say where a claim came from.
Most of what Antibody’s agents did was approved by nobody. Cleanups and self-corrections ran under policy: this class of action, in this system, within ±15%, with a re-check. For those, the record carries the action and its operation type, and nothing about the policy that permitted it or the envelope it was held to. In a system that heals itself, policy authority is the common case, and the record does not yet express it.
Entry 7: What the evidence showed us
The record’s most useful contribution was not flagging the valve. We had declared the valve consequential ourselves, so the flag restated our own decision. The record’s value was that reading it showed us where the system’s account of itself was inconsistent: the valve typed as a read, the hosted run with no profile, the proposal recorded as an execution.
That value has a boundary. The record contains what the application routes to it, decided by the same process it describes. It does not observe the building, record payloads, or resist tampering. It is structured execution evidence of what each agent reported doing against what it declared: materially more than an application reporting “success,” and not an audit trail or a flight recorder.
What Antibody changed for Governor
Governor itself did not change. The integration ran on the released Sentience Governor 0.3.2.1, and everything above was built in the application. What changed is what we now know Governor needs.
Governor’s core primitives traveled into another runtime and another domain: an agent with an objective and a scope, actions evaluated against them, a record per agent. Four assumptions around those primitives did not travel, and each points to a design input. None of these are shipped capabilities.
- Meaning must be declared. Operation types should be declared by the operator, by tool pattern in the profile, rather than inferred from names. This matters most for the planned Level 1 enforcement, where an action typed as a read would never be held.
- Policy must accompany execution. Profiles have to reach wherever the agent runs, and a session with no profile has to look different from a clean one. That includes policy authority: the record should say which rule and envelope permitted an autonomous action, not only which person approved a held one.
- Authorization belongs to an action lifecycle. Proposed, held, authorized by whom, executed and result should be first-class states, with the source of each authorization claim. We expect to design this inside the planned Level 1 prompt mode, which produces exactly these states. A public synchronous session lifecycle belongs here too, so threaded runtimes need no private calls.
- Evidence needs a path out of the executor. Remote execution needs a supported way to bring records home. Whether governance evidence can travel across a distributed agent system without centralizing the underlying execution data is the open question Flower leaves us with. We exported one executor’s summaries to one operator. We did not build federated evidence.
Two problems remain open without a design yet: the record cannot show whether an action achieved its effect, and nothing in it is tamper-evident.
The test we would apply to any system now
Pick the most consequential tool your agents can call. Ask your governance layer what kind of operation it believes that tool is, and where that belief came from. Then run the agent somewhere other than your own machine and check whether the answer, and the profile behind it, came along.
We kept Antibody running after the hackathon and published it, with a few enhancements: a Governor view of the records across runs, and approvals tied to a verified session. Try it at demo.getsentience.ai/antibody.