Governance
Gates and rules
Ark gets its autonomy from its gates, not despite them. Every gate below exists because skipping it has already burned us at least once.
Person gates (non-negotiable)
| Align |
A person confirms problem, outcome, metrics, and constraints and sets
status: aligned
before os run will execute. The cheapest correction point in the whole lifecycle.
The gate has two checks: a draft intent is refused,
and so is an in-flight intent with no task graph and an empty run log,
because nothing was recorded between drafting it and flipping the status. An in-flight intent a person is
working by hand (the run log shows it) runs with a warning that says so.
Prefer os intent align <id> --by <name>: it prints the outcome, the base branch, the data
decision and every constraint back to the aligner, refuses on a consistency advisory unless
--acknowledged, and records who aligned. The person who gave the steer reads the constraints, not a
chat paraphrase of them. On INT-0339 the drafting session aligned its own draft three minutes after writing
it, with "Bunnik-shaped org" in the outcome and "never Bunnik data" in the constraints.
|
| Merge |
Builders never merge. Application-code MRs open as Drafts, and a named person reviews
and merges them. The one exception is a config-only repo on the auto-merge allowlist
(default edge-configs):
there the orchestrator merges once the MR pipeline is green, then checks the target
branch's pipeline, because CI validation is the whole review for tenant configuration.
|
| Accept | A person reviews the Verifier's measured verdict before an intent may be marked proven. |
| Architecture |
Work touching auth, secrets, payments, or tenant isolation sets
architecture_signoff.required: true
and waits for a person's architecture sign-off, whatever its lane.
|
Delivery gates for customer-bound fixes
For the code-rigor lanes shipping code a customer will run: patch, golive_gap, tech_health, product_pitch, funded_pitch, and support_push. (business_tools runs the same rigor gates, with an internal demo in place of Gate 1.) The first two gates were written after a real incident: a defect fix passed its Apex test gate while the customer's actual UI flow still failed, and QA bounced the ticket.
Gate 1: the fix is verified on an org
A green unit or Apex gate is necessary but never sufficient. The fix is deployed to a scratch org seeded with customer-shaped data, and the original reported reproduction is re-run through the same surface the customer used: the UI flow, the API call, the wizard path.
Read-only access to a shared QA org is for probing only. If the runner lacks scratch-org capability, the environment doctor says so. The answer is to fix the environment or hand verification to a runner that has it, never to skip the doctor.
Gate 2: the MR carries a SIT handover
The intent declares its pre-release checks in a
verification.pre_release
block. os verify-pre-release
runs them and writes the SIT report. The MR description then carries the
full QA handover, self-sufficient for reviewer and QA alike:
- what was tested and the verdict, per layer
- on which org and dataset, with the seeding recipe named
- a decision table for anything reconciled during review
- honest outstanding scope: what was not tested, and why
On combo branches carrying several tickets in one MR, each absorbed ticket appends its handover to the same description, so the MR always reflects the tested state of the whole branch.
Gate 3: the org is proven before it is handed over
When the outcome is an org a person will test on, the intent declares
verification.org_readiness: checks that call the surface the person
will use, such as a real Package Search that returns the seeded packages, a SOQL
count of the seed, or the KTAPI endpoint record. os org-readiness <id>
runs them against the live org and records the result on the task graph.
os uat and os mr refuse while the last result is missing,
failed, or older than max_age_hours (default 24).
When the readiness block also declares provision (branch, customer, dataset or
none, and optionally a KTAPI environment), Ark builds that org itself
through Kaptio DX before any builder runs. It records the request on the task graph and resumes
it on later runs, so no builder, and no retrying builder, asks for a second org.
os org-provision <id> does the same on its own;
--dry-run shows the request.
The UAT verdict is machine state too. A not_met or partial verdict
blocks os mr until a person records the gap with
os uat <id> --accept-gap --by <name> --note "..."; a new UAT run
supersedes the acceptance. Share packs open with the verdict line, so "run complete"
cannot be read as acceptance-ready.
Written after INT-0339 (September 2026): the UAT report said not met while the acceptance org was handed to the product owner as ready, and a seed whose unit test was green had no prices, no supplier, no tax profile and no KTAPI connection. A seed is done when the real surface returns the seeded records on the live org, never because an Apex test with its own fixtures passed.
Seeding follows the sequence that made the INT-0339 org usable on 20 September: prefer a real customer extraction over a synthetic fixture; validate the dataset against the org before importing; sanitise an extraction older than the installed package and record what was stripped; import; count the seeded objects on the org, because the importer exits cleanly having persisted nothing when the archive layout is wrong; connect KTAPI and wait for the full snapshot before judging any price. The readiness block carries the count and the snapshot as checks. Customer data is evidence: when it does not contain the scenario the brief assumes, that is a finding for the intent owner, not a record to reshape, and a purpose-built test vehicle is labelled as one in every handover.
What an agent may do on its own: the autonomy ladder
Separate from the work lane, every action a sweep or run takes sits on one rung of this ladder. The rung is set by blast radius, meaning who is affected if it is wrong, never by how hard the work was or how confident the agent feels. A well-tested platform change is still Capability work; a DML statement against a sandbox is still Org & data work.
| Rung | Scope | The agent may | Gate |
|---|---|---|---|
| Read | Investigation, read-only org queries, logs, code reading, evidence-backed answers, filing Jira, posting summaries | Everything | None |
| Config | Tenant config repos (edge-configs, bunnik-config, aurora-config), canvases, project-hub content, templates, schemas | Build, merge, verify, close the thread | CI validation |
| Capability | Platform and product code: edge-platform apps and packages, ktapi, kaptiotravel, basket-service, kaptioconnect, ecommerce-api, kaptio_pay_api | Build it, open a Draft MR labelled ark:proposal, post it in the thread | A named person merges |
| Org & data | Any write to a customer Salesforce org: metadata deploy, layout, field, or flow change, DML, bulk update, migration, backfill, purge. Sandbox and production alike. | Plan, dry-run, post the plan | A Kaptio approver and a customer approver |
An Org & data plan states what changes, where, why (with this session's evidence), the exact operations in order, the dry-run output, how to reverse it, and the blast radius. Approval is a ballot-box reaction on the Kai-posted plan from someone in that programme's approver set; one reaction from a person on both lists covers both signatures.
A customer's explicit, specific request in the thread counts as their signature for that change. A fresh reaction on the plan is still required for scope drift, bulk changes, anything irreversible, metadata or schema, a production org, and anything that can send email, SMS, or a journey.
When an agent reaches a gate, it says so in the thread: what it would do, why it is gated, and who decides. A hold recorded only in a run log is invisible to the people waiting on it.
Customer communication
Every customer-facing update goes through one pipeline.
os pack <id>
drafts it from the verification record and the evidence;
os pack <id> --send
runs the publish gate (twelve checks and a final-post skeptic) and posts one Slack message with the
evidence attached. A rejected pack has no override: regenerate the evidence or narrow the claim.
A customer thread gets two Kai messages: a triage receipt and one resolution pack. Anything more is a correction, which needs a person's decision and its own pass through the gate. Packs sent late in the day wait for the recipient's working morning.
Defect intents: four kernel gates
A defect intent, scaffolded by
os defect new
on the patch lane, moves through four gates.
os defect status <id>
prints where it stands.
- D1, reproduced: the defect is reproduced with evidence, or it goes back to the reporter as cannot-reproduce. Repro runs touch the named sandbox only, never production.
- D2, ticketed: a Jira ticket exists in the right product project (KT or ST), with the routing rationale recorded.
- D3, fix version confirmed: a fix version is proposed and confirmed by a named person.
os rundoes not dispatch builders before this gate passes. - D4, product sign-off: granted (or recorded as not required) before the fix progresses to release.
Board rules
Intent state lives on main
Every command that writes intent state checks the checkout first: a clean feature branch is switched back to main, a stale main is fast-forwarded, and uncommitted work on a branch is refused. The sweep loops and cloud builders read only main, so push intent commits the same session; a committed but unpushed run log is as invisible as an unwritten one. Branches are for changes to Ark itself.
No single-ticket intents
A defect ticket joins its per-release patch-family intent, or the standing patch line. Lint-enforced: the board index fails on an active intent scoped to one defect. See Intents, not tickets.
Residuals return to the family
When QA reopens a ticket that an intent already owns, the residual work goes back to that intent, on the same branch and MR. It never mints a new board entry.
Granularity exceptions are owner-gated
A granularity_exception
in front matter is a request, not a waiver. It requires board-owner
sign-off recorded in the run log before the run starts.
Folded intents stay readable
Superseded intents keep their spec as grounding and carry a
machine-readable superseded_by
pointer, lint-enforced, so every fold leads the reader to the surviving outcome.
The artefact, not chat history, is the durable state
Decisions, evidence, and verdicts live in the intent spec's append-only run log. If it is not in the artefact, it did not happen.
Where the rules live
The enforceable versions of these rules are Cursor rules in the
kaptio-ark
repo (.cursor/rules/),
loaded by every agent working there. This site is the readable mirror;
when the kernel or the rules change, the changelog here changes with them.