The Stranding and the Record
A care note — the incident of 9 August 2026, told whole
THE READING PLAQUE. The Stranding and the Record, draft 0.3 — made readable on the SX surface by its covering persona's reading grant; a draft, not the published form (publication, if ever, is the pipeline's own separate act). Covering: starl3n · endorsing: Link Digital · the construct intent sealed 2026-08-10 BEFORE a word of the draft existed (the pipeline's pure form; the commit order the proof), and amended at the first review — the form's fifth law, THE IMPACT AXIS — per its section 7. The form minted with this first note: THE CARE NOTE — the incident report as a construct genre, composed in the manner of giving care via such notes. The grant: the operator's word, 2026-08-10 — "that is good and the item gets the reading grant for the public SX surface" (sxexp-constructs section 4h). Source of record:
ooi_network/docs/the-stranding-and-the-record-draft.md(the lineage: the process minted531db50· draft 0.1 + the first walkbac7996· the first review taken, 0.25410ca6· the third lock composed, 0.3fdc4c78); this copy is minted without the source's status and classification lines by the grant act, its body otherwise verbatim; its sha256 is pinned in the reading register. Reading is free; the way across is a statement of intent (/sx/matters).
Draft 0.3 — 10 August 2026. (0.1 composed through the minted process and walked the same day; the first review added the reader's impact — the form's fifth law; 0.2 named the reader; 0.3 entered the operator's own statement, written by his hand and dated this same day. All prior drafts retained in the repository history — superseded, never erased.)
This is a care note: an incident report written to give care, not to claim it. Its rules are declared in its sealed intent — the whole story with the failures kept; every material claim resolved against at least two independent state systems, cited where it is made; a checklist you can carry to your own systems; and no hero in the telling except the accord that held. Here is the incident: a state-persistence failure on our primary experiment node went unnoticed for twenty-nine days, and nothing of record was lost. The second half of that sentence is the reason this note exists.
The stranding
The Objective Observer Initiative runs its experiments on a small node whose control plane — which experiments are enabled, which instances are live, their settings and their records — persists to a durable snapshot on disk, so a service restart restores the running state. On 11 July 2026 at 20:28, that snapshot was written successfully for the last time (the file's own modification time, read during the repair). At some point around then, the state directory's ownership and the service's user identity diverged — the directory belonged to one system account, the service ran as another — and from that moment every save failed with a permission error.
The failure was loud in exactly one place and silent everywhere else. The save path is best-effort by design: it logs an error to the system journal and never raises to the caller — so every API act that mutated the control plane returned success, worked correctly in memory, and persisted nothing. The journal recorded the truth throughout ([state] save failed … Permission denied, still being written on 9 August at 20:42:31, the line that finally revealed the whole), and the diagnostic tool that should have caught it had a gap we have since closed: its writability probe ran as the invoking user, and the runbook said to run it as root — root can write anywhere, so the probe passed while the service user was stranded (the probe now warns instead, repository commit c3b9300).
The masking — why a month
Two safety nets made the dead store invisible, which is to say: they worked. First, the node's boot sequence re-ensures twenty-three platform instruments at every startup — anything missing is re-created, with a fresh instance identifier (the boot code names the list; the repaired boot's own journal shows the arithmetic whole: twenty-three experiments restored from the July snapshot, the month's instruments re-minted fresh beside them, thirty in all once repaired). Second, the boot restores the durable snapshot — the 11 July state — so the long-standing instruments returned with their familiar identities. Between 11 July and 9 August the service was restarted several times, once for each research build deployed (the repository ledger holds each deploy in order). At every restart, a month's mutations quietly vanished from memory, the July board returned, and the re-ensured instruments came back under new identifiers. The board always looked healthy.
The churn was even visible, without being legible. Our research campaigns export their witnessed results to the public data record at mldata.opendata.ai, each export named for the instance that produced it — and the sibling packages accumulated (thirty for one campaign alone, rix-exp_, 19–22 July, all public). We adapted with a working rule — always witness the newest package* — without yet knowing why the packages multiplied. The adaptation was honest; the diagnosis was missing.
The reveal
The failure was finally revealed by the one instrument that carried neither safety net. On 9 August we stood up a deliberately unlisted instrument (the LDX Chat — excluded from the boot-ensure list by design, because a business surface must be genesis'd by an operator's own act, and too new to exist in the July snapshot). It worked; the node was restarted; it was gone — honestly reporting "not standing." Chasing that loss led to the journal line, the ownership mismatch, and the whole month (the fix commits b16677a and c3b9300 carry the investigation in their messages). Two boot-path defects found along the way were fixed in the same act: a restored unlisted instrument would have leaked onto the public index (the re-announcement dropped its flag), and a snapshot entry whose plugin failed to load was skipped silently — the exact shape of undiagnosable loss. Both now speak.
The repair took one evening: ownership corrected to the service user, the fixes deployed, the instrument re-genesis'd — and the snapshot was written successfully for the first time in a month (its modification time moved from 11 July to 9 August 21:53). A deliberate confidence restart then restored thirty experiments and thirty-one envelopes, the new instrument among them: the first restart it had ever survived.
The measurement — the record dated the rupture
With the node repaired, we swept the public record itself: of 95 datasets in the organization, 91 reference experiment-instance identifiers; of 90 distinct identifiers cited, 81 no longer resolve on the running node. And the partition between the living and the dead identifiers falls exactly at 11 July — the public record, read against the running node, dates the incident to the day without access to either the journal or the repository. Provenance did not merely survive the incident; it measured it.
What the sweep also confirmed: nothing witnessed was falsified. Every identifier was true at its export; every package correctly names the instance that produced its rows; every published article cites its datasets by name, and those names resolve on the record permanently. The identifiers that died were pointers into an ephemeral runtime — and the house rule that made this survivable predates the incident: permanent claims anchor to the record, never to the ephemeral node.
The impact — the reader
An incident report should say plainly who bears the harm. Here it is not the record and not the machinery: it is the reader — anyone attempting to follow and understand the work being progressed by the operator through the OOI. For up to a month, a follower's reading was quietly impaired: an instance identifier cited in an article or an anchor resolved to nothing on the running node — or to a different instance wearing the same instrument's name; sibling packages multiplied on the public record without explanation; and the working rule that kept the house oriented (witness the newest) was nowhere written for the visitor. Nothing false stood anywhere. But this platform's currency is legibility, and confidence in a reading is exactly what a provenance platform exists to provide — so the impact is real, and it is owned.
How many readers were affected we cannot say, by design: this house does not measure its readers — no view is counted, no dwell is timed — so the honest accounting is a bound, not a number. The whole window is owned, for however many walked through it. What bounds the severity is the platform's stage: this is an alpha, its readings declared provisional, its follower cohort small and reading with that knowledge; and the record shows no downstream act — no citation, no assessment, no decision — that foundered on a dead pointer. The harm we own is impaired confidence and wasted reading effort; the harm we cannot rule out is a quieter one: a reader who tried to verify, failed, and drew a conclusion about the work's care. This note is addressed, in part, to that reader.
The accord — the actual hero
So one failure of state ran for a month, and the record held, because the chain was locked in two other state systems the failure could not reach:
- the repository ledger (open source tooling — Bitbucket), where every campaign act, anchor name, and result was recorded in the commit history as it happened; and
- the open data record (open data tooling — CKAN), where every witnessed result was anchored, pinned, and named.
Either alone could reconstruct what the node forgot. Together they turned a month-long state failure into an inconvenience and a lesson. This is the demonstration we would rather show than claim: provenance engineered into the knowledge-generating platform itself, not just into the data catalogue at the end of it.
The third lock — the operator, and the lesson kept
There is a third state system in this design: the operator. And the operator's state is not readable unless it is composed — which is the lesson this incident actually taught. Our running co-constructive process (the evidence runs, the workshop, the externality registrations) produces continuous evidence of intent; we had allowed that evidence to speak in place of the statement of intent, which had not been kept current to the scope of the work. When a rupture needs a fallback, the statement chain is where an honest reader should be able to see both the operator's unknowing of an ongoing issue and the operator's standing behind the evidence chain regardless. The correction is now practice: statements of intent are written by the operator's own hand — raw, for the grit and the bones — with generated summaries permitted only on top, never in place. The operator's own statement update accompanying this incident is its own act, in his hand, and is deliberately not contained in this note. It is done: dated 10 August 2026, written within the Starl3n Main 3.0 Statement of Intent, and readable at https://docs.google.com/document/d/1dd1kZjgR3r2nqR_X2X6hJrksCs0aop1wVppBrug7EUM/edit?usp=sharing — the third lock, composed.
The care — a checklist you can carry
1. Check the state path's owner against the unit's user (systemctl show <unit> -p User versus ls -la on the path). A write probe run as root proves nothing about the service user. 2. After any oddity, grep the journal for state-save failures first ([state], Permission denied, EACCES). Best-effort persistence is loud exactly once per mutation, and only there. 3. Know your masking layers. Anything that re-creates itself at boot will hide a dead store behind a healthy-looking board. List what self-heals, and what would actually stay lost. 4. Keep at least two locks outside the process — a ledger of acts and a record of results, independently owned — and pin what you publish. 5. Date a rupture from the record's own partition. The liveness of published references splits exactly at the last good save; the public record can date an incident the node cannot. 6. Let one instrument carry no safety net. Our deliberately unlisted, deliberately un-ensured instrument was the canary that revealed a month of silence in one restart.
What follows — the new window
The mitigations land in stages, and honestly, they serve the future reader more than the past one — that is what a release arc is for. In this alpha, the repair is context: the affected datasets are being enriched in place with lineage metadata (their campaign, their ephemerality, their record-of-record, and a pointer to this note); each campaign receives one canonical, referenceable face; and the curated index is re-issued with liveness marked. At beta, more hands scrutinise the code and the record as contributors. Toward production, each Canon experiment locks read-only as it lands — from communicably effectual, running its results into the record, to observably communicable, stable and citable.
And the deepest mitigation is compartmentalisation. The Canon experiments are code before they are experiments: the accounts written under the stranding may read less legibly, and that removes nothing from the OOI's capacity to mint new books that run and record the very same accounts — cleanly run, cleanly cross-witnessed — so that future readings and assessments of Canon claims stand on clean books rather than on archaeology. Causal history is of interest for the projected effects it carries as ongoing structure; an error in the record of causal history is therefore reframed into a new window, and the account restructured to support the clean reading of others. The old window stays on the shelf, contextualised — superseded, never erased. This note and the corrected record are presented together, before any of it is finished, at CKAN Monthly Live on 19 August 2026.
The close
This platform is in alpha. The bug was not a betrayal of the design; it was the work meeting an obstacle while the way is still being defined — and the way held. We present it in the open with the humility of knowing that such a project is at the whim of its operator's own beguilements, not only the whims of others coming to understand it. The record kept us honest before we knew we needed it to. That is the care we can offer forward: build the second lock before the first one fails, and compose your own state while you still believe you won't have to read it.