Referentautomat
The org-owned GitHub App that handles PR hygiene and the release handover, covering area labels, Linear attachment, RC title validation, the pre-RC audit, and the drafted release notes.
Referentautomat is the GitHub App that does the repetitive parts of our PR and release process. It labels pull requests from the paths they touch, keeps them attached to their Linear issues, audits a Release Candidate before it ships, and drafts the release notes.
Every write it makes carries its own identity (referentautomat[bot]), not
github-actions[bot] and not a maintainer's personal token. That is the point
of it being an App: the audit trail names the automation, and the releases page
stops being attributed to whoever happened to own the token.
Authority Model
The bot's authority is split by determinism, and the split is deliberate.
- It acts on its own where the same input always yields the same output. Area labels from a committed path map, RC title validation, wrong-base warnings, the readiness audit. These are mechanical, so there is nothing for a human to approve.
- It only proposes where judgment is involved. A PR with no ticket gets a comment listing candidate issues, never an auto-created issue. The measured failure mode is an unlinked existing ticket, and auto-creating would duplicate it.
The one place machine-written text ships to customers is the release notes region of the RC body, which publishes as the GitHub Release and the Linear changelog. The reasoning behind that, and the alternatives rejected along the way, are in ADR 0001.
Where It Lives
| File | What it is |
|---|---|
.github/workflows/referenta-bot-pr.yml | Runs the PR checks on every pull request. |
.github/workflows/referenta-bot-release.yml | The RC audit and the release-notes draft. |
.github/workflows/referenta-bot-tests.yml | The bot's own unit tests. |
.github/referenta-bot/scripts/ | Plain Node ESM, no build step, no dependencies. |
.github/referenta-bot/label-map.json | The path-to-label map. |
.github/referenta-bot/skills/release-notes/SKILL.md | The house style for the drafted notes. |
.github/referenta-bot/skills/qa-checklist/SKILL.md | The house style for the drafted QA test steps. |
.github/referenta-bot/qa/parent-header.md, qa/surfaces.json | The QA checklist's parent header and the product surfaces with their staging routes. |
.github/workflows/referenta-bot-issues.yml | Triage and fix on GitHub Issues. |
.github/referenta-bot/issue-policy.json | Denied paths, budget, fixable kinds, models per stage. |
.github/referenta-bot/skills/issue-triage/SKILL.md, issue-fix/SKILL.md | House rules for the two issue stages. |
The release.yml workflow also mints an app token, so the git tag and the GitHub
Release are attributed to the bot. See Continuous Integration.
On Every Pull Request
referenta-bot-pr.yml runs on opened, synchronize, reopened, and
edited. The edited trigger is what makes the checks self-clearing: adding
Ref REF-1234 to a description re-runs them and the finding disappears.
All findings land in one sticky comment that is edited in place and deleted outright once nothing is wrong, so the bot never accumulates comments on a PR.
The workflow run itself appears in the PR's check list and passes whether or not it found anything: a finding is reported in the comment, and the job fails only when the bot itself breaks (a bad token, a GitHub API error). Nothing it reports is a required check, so nothing it reports blocks a merge. The comment is the signal; the check is not.
Area Labels
Labels come from label-map.json, matched against the PR's changed files. A
file counts toward a label when it matches any include glob and no exclude
glob.
{
"label": "Assistant",
"include": [
"apps/dashboard/app/[lang]/(main)/assistant/**",
"apps/dashboard/app/[lang]/api/chat/**",
"apps/dashboard/lib/assistant/**",
"apps/dashboard/app/tools/**"
]
}The glob syntax is deliberately minimal: ** spans
directories, * stays inside one path segment, and every other character is
literal. That is what makes the dashboard's route groups safe to write
directly, since a general glob engine would read [lang] as a character class
and (main) as a group.
Labelling is add-only. The bot never removes an area label, so a label you added by hand survives, and a PR that outgrows its original scope accumulates labels rather than having them churn.
Adding a label to the map only works if the label already exists in the repo. The GitHub API creates missing labels silently with a random colour, which is how you end up with two near-identical labels. Create it in Issues ▸ Labels first, then add the map entry.
Linear Attachment
Linear attaches a PR to an issue when the identifier is in the branch name or
the PR title, or when a magic word precedes it in the description. A bare
REF-1234 in the description alone does nothing, which is the failure mode
this check exists for.
The bot handles three cases:
-
Already attached. Nothing to do. It also removes the
needs-linearlabel if one is left over. -
Mentioned but not attached. It appends the magic word to the description for you:
<!-- referentautomat:linear --> Ref REF-1234 https://linear.app/referenta/issue/REF-1234Editing the description is what re-fires Linear's webhook, so the attachment appears immediately.
Refis a non-closing magic word, so merging the PR does not complete the issue. -
No reference at all. It applies
needs-linearand comments with up to three candidate issues, found by searching Linear for words from the branch name and title and keeping only those whose own titles share real words with them. Candidates are suggestions to click, never applied automatically.
Three guardrails on case 2. Identifiers inside code spans and fenced blocks are
ignored, so a PR documenting this syntax does not read as attached to its own
example. A PR that already has an attachment is never touched, so a
cross-reference like "supersedes REF-900" is left alone. And the block is
written at most once: the marker comment is the bot's record that it has acted,
so deleting the Ref line while leaving the marker is a durable way to reject
the guess.
It does write for every bare identifier it finds on an unattached PR, so a ticketless PR whose description discusses other issues in prose will get them all linked. Attaching the right ticket first is what prevents that.
PRs from other bots are skipped entirely. Dependabot and friends legitimately
have no ticket, and flagging them would train everyone to ignore the comment.
Referentautomat's own PRs are the exception: they carry the issue's REF-N
when it named one, and are flagged like a human PR when it did not.
Wrong Base Branch
A PR that targets main without coming from dev gets a comment telling you to
retarget it, because the only PR into main is the Release Candidate. It is a
warning, not a block: the bot does not retarget the PR for you. See
Git Workflow.
Bot Pull Requests
A PR authored by referentautomat[bot] gets two extra checks. Its current
diff is re-checked against issue-policy.json, so a push that later crosses a
denied path or the budget is reported in the sticky comment. And if the actor
who marked it ready for review is the bot itself, it is converted back to
draft: only a human promotes a bot PR. Unlike other bots, Referentautomat's
PRs are not exempt from the Linear attachment check.
On The Release Candidate PR
A PR whose head is dev is the Release Candidate by definition, and it is
treated differently: area labels and ticket checks there would just be noise,
since the RC contains every PR in the range. What it gets instead:
Title Validation
The title must match <Prefix>: vX.Y.Z exactly. A near-miss (missing v,
missing space, trailing text) makes the version tag, the GitHub Release, and
the Linear sync all skip silently at merge, so the bot reports it while the
PR is fresh and suggests the corrected title. A title that is not
release-shaped at all is reported too, since a dev to main PR is an RC
whether or not it is titled like one.
Codesmith Footer Stripping
The RC body is the public release notes and gets mirrored into Linear, so the
vendor footer does not belong there. Feature PRs are deliberately left alone:
codesmith re-appends the footer in response to the edit, which fires the
edited trigger again, and the two bots take turns rewriting the description
(observed for eight rounds on one PR before this was scoped to the RC).
Pre-RC Readiness Audit
The audit walks main...dev, maps every commit to the PR that carried it, and
posts one sticky comment covering:
- merged PRs with no Linear attachment,
- PRs missing the area labels their own diff would earn,
- PRs adding a
supabase/migrations/file with no matchingsupabase/schemas/change, which is the zero-drift rule, - commits that reached
devwith no PR at all. - how far the release notes cover
dev: the PRs merged since the notes were last drafted, or the reason no draft exists yet.
A clean range says so explicitly rather than staying silent, so you can tell
"audited and fine" from "never ran". When the notes are current, the header
names the dev commit they cover. The audit is purely informational and
identical in dry run and live: it reads, reports, and changes nothing. On an
event that also drafts notes, the audit waits for the draft and reports the
result, so its coverage line is never one draft behind.
Release Notes
The bot drafts the notes into a region of the PR body delimited by
<!-- referentautomat:notes --> markers, so the prose around the region is
yours and survives every refresh. The first line inside the region is the
notes head, an HTML comment naming the dev commit the notes were drafted
for. It is how the bot knows whether the notes are behind dev, and every
renderer that hides the markers hides it too.
The draft itself is one chat request through the Vercel AI Gateway, so spend
and traces land alongside the dashboard's. The model is set in
.github/referenta-bot/release-policy.json: a DeepSeek model, picked by
replaying a real RC against Claude. It gets no tools. The release context, the
classifier's verdicts and a trimmed diff per PR travel in the message, with
SKILL.md as its house style: bilingual with German first, written in plain
language for policy and public-sector staff with no technical background
(effects, not mechanisms; no technical terms or layout details; at most four
one-sentence bullets per section), internal work collapsed into a single
catch-all line for users, flag-gated work carrying an availability note. A delta of internal PRs only is appended without a model
call. The reply becomes NOTES.md, and a separate deterministic step splices
that into the body. The job summary lists the request's tokens, reasoning
tokens included, and the gateway's cost.
Internal work is collapsed into one catch-all line for users, and listed for
developers in a region between <!-- referentautomat:internal --> and
<!-- /referentautomat:internal --> at the end of the notes: one English bullet
per internal PR with its link. public-body.mjs keeps that marker pair (GitHub
hides comments, and the dashboard strips the whole region in
lib/github-releases.ts, so the in-app page never shows it); the Linear mirror
in linear-release.yml runs the script with DROP_INTERNAL_MARKERS=1, so the
section shows in Linear without comment text.
Everything else in the bot is deterministic Node. The model is used for exactly one thing, which is turning a commit range into customer-facing prose.
Prose already in the region is treated as authoritative. A refresh adds what is missing and corrects what the code now contradicts; it does not reword what you wrote. Polish the German and English freely, then refresh without losing your edits.
This extends to a hand-written body with no markers at all: it is adopted as the existing notes rather than being overwritten.
Finally the bot applies Release Candidate and sync-to-linear, in that order
and only after the notes exist, so the first Linear sync mirrors a body that
already has content instead of creating a release with empty notes. Full release
flow in Release Automation.
When The Notes Refresh
The draft runs when the RC PR is opened or reopened, and again on every
push into the RC whose head differs from the notes head, which is to say
whenever something new merges into dev while the RC is open. Each refresh
revises in place: a docs-only merge typically leaves the prose untouched and
only advances the notes head. A push that does not move past the notes head
(a re-run at the same commit) drafts nothing. The audit re-runs on every push
regardless.
Nobody needs to open the RC by hand. The Release Candidate workflow
(referenta-bot-release.yml, Actions → Run workflow) opens it as a draft PR
titled with the next version (ensure-rc.mjs), and that PR's opened event
runs the first draft like any other. The dispatch drafts nothing itself. GitHub
offers Run workflow only for a workflow_dispatch workflow that is on the
default branch. Inputs are listed in
Release Automation.
QA Checklist
The bot keeps one Linear issue per Release Candidate that carries the
QA Required label, titled QA: vX.Y.Z, with one sub-issue per product
surface, holding the manual test checklist QA works through on staging. Not
every RC gets a manual pass, so the label is the opt-in: an RC without it is
left alone, and adding the label to an open RC creates the checklist right away
(the workflow listens to labeled; every other label resolves to nothing).
Removing the label stops further refreshes but leaves the existing issue in
place. The parent is created in the team named by LINEAR_QA_TEAM_KEY, labelled QA Checklist, placed in the project
named by LINEAR_QA_PROJECT_ID when one is configured, attached to the RC PR
(attachment Release Candidate), and added to the Linear release for the
version when one exists. Sub-issues share team, label, project and release; a
run adds the release to any checklist issue still outside it, so a release
created after the checklist is joined on the next push. The bot never changes a
status or assignee: sign-off stays a human step.
The qa job runs after notes in referenta-bot-release.yml, since both
write to the RC PR body and running them in parallel could lose one's update.
It still runs when notes was skipped or failed, as long as resolve
succeeded.
The checklist is all plain Markdown so Linear renders the checkboxes natively:
- Parent
QA: vX.Y.Z: a committed header (qa/parent-header.md: staging URL, QA account, the feature-flag rule) and aRelease notes:link to the Linear release notes of the version, added as soon as the release exists. Nothing else: Linear's sub-issue list is the overview. Anything a human adds under a heading of their own is kept. PRs with no user-visible effect get no section anywhere; their numbers are recorded in the RC body stamp (below) so they are not drafted again. - Sub-issues
QA: vX.Y.Z · <Surface>: one per surface inqa/surfaces.jsonthat has a testable PR, opening withOpen: <staging route>, then one### <PR title>section per PR with a Where line (surface, ticket, PR URL), an optional deeper Open URL, and three to six checkbox steps drafted by the QA model from the PR body and diff, followingskills/qa-checklist/SKILL.md. Headings carry titles, not PR numbers: QA has no GitHub access. The PR URL in the meta line is the key the merge uses.
Feature flags are on in staging (dev.referenta.de is a preview environment
for apps/dashboard/flags.ts) and off in production until activated, so a
flagged change gets a first step saying exactly that and is otherwise tested
like any other.
Adding a surface means adding it to qa/surfaces.json; the draft step is given
the list as QA_SURFACES.md and the apply step refuses a section whose surface
is outside it.
A refresh only ever appends. Merged PRs are immutable, so a fix always
arrives as a new PR. On every push into a labelled RC whose head moved past the
checklist's, the bot asks the model for sections for the missing PRs only and
merges them in deterministically. Every existing section is kept byte-for-byte,
QA's ticks and notes included. The only line the bot rewrites is a section's
meta line: the PR becomes a #<number> link and the PR author is named, so QA
can ask the right developer without GitHub access. A PR that left the range (a
revert) has its heading suffixed (reverted, skip), and loses the suffix again
if it comes back; a reverted PR without a section leaves the stamp's skip list. Only PRs
that are actually merged count as in the range: GitHub also associates a dev
commit with any open PR whose branch has merged dev, and those are ignored.
Preservation is enforced by code (qa-sections.mjs), not by the prompt.
A checklist created before the sub-issue layout (its sections are headed
### #<number>) is rebuilt once on the next run: the parent keeps its identity
and release link, its per-PR sections are re-drafted into sub-issues, and ticks
on the old sections are lost. A parent that still carries a ### Core smoke test, ### Internal changes (no manual test) or ### Changes section from
an earlier sub-issue layout loses them on the next run; the PRs those sections
listed as internal move into the stamp's skip list.
The bot's bookkeeping lives at the top of the RC PR body:
<!-- referentautomat:qa issue=<linear-issue-id> head=<dev-sha> skip=<pr>,<pr> -->
**QA checklist:** [QA-171 · QA: v3.8.0](https://linear.app/referenta/issue/QA-171)
<!-- /referentautomat:qa -->The markers are hidden on GitHub; the middle line is a quick link for
developers. This block never reaches the published notes.
scripts/public-body.mjs removes it, along with the notes markers and the
Codesmith footer, and both the GitHub Release step in release.yml and the Linear
mirror in linear-release.yml build the public body through it. The dashboard
reads the GitHub Release, so it is covered too.
If the block is lost (someone rewrote the body), the next run finds the open issue by its exact title and re-stamps the body rather than creating a duplicate; the skip list starts over, so the PRs without a manual test are drafted once more and marked internal again. The same fallback covers a stamp that points at an issue which no longer exists. A Linear API failure fails the job: unlike candidate search, the write is the point.
A @referentautomat qa comment lands in its own concurrency group, so it can
run concurrently with a push-triggered run on a brand-new RC; if that race
creates two QA: vX.Y.Z issues, delete the duplicate by hand — later runs
converge on the remaining one.
Commands
Comment on the RC PR to trigger either capability by hand:
| Comment | Effect |
|---|---|
@referentautomat notes | Refresh the notes region, revising in place. |
@referentautomat audit | Re-run the readiness audit. |
@referentautomat qa | Refresh the QA checklist: draft sections for merged PRs that have none. Works with or without the QA Required label. |
The command must start a line of the comment, and the commenter must be an
owner or member of the organization, or a collaborator on the repo (GitHub's
author association). A mention in passing, or a command
from anyone else, is logged and ignored. notes and qa additionally require a
valid RC title, since without one the release pipeline would skip anyway and
drafting notes or a checklist for it would be misleading. The audit has no such
requirement and runs on any RC.
A command works only once the workflow that knows it is on main. GitHub runs
issue_comment workflows from the default branch, not from the PR, so a
command added on dev does nothing when commented on the RC until that RC
ships; the run still shows as successful because the old workflow finds no
command to match. Until then, trigger the job by pushing to dev (a merge into
the RC synchronizes it) or by re-running the failed job from the Actions tab.
Comment bodies are untrusted input. The decision of what a comment asks for is
made in release-policy.mjs, which only ever matches the text; nothing from a
comment is interpolated into a shell or a prompt. Each command runs in its own
concurrency group, so one queued behind a running audit is never cancelled in
its favour.
On Issues
Since 2026-09 the bot also works the GitHub Issues queue. Everything in Linear
stays in Linear; GitHub Issues is where a bug or chore goes when the bot may
take it. File one with the Bug or Chore form; the fields are what the
bot needs to act. Name the REF-N in the form so the fix PR attaches to it.
Triage
Every opened or edited issue is triaged by Claude Code (Sonnet 5) and reported
in one sticky comment: the kind (bug, chore, feature, question), one
area label, what is missing (needs-info), likely duplicates (duplicate),
and Linear candidates when no REF-N is named. Kind and area labels are
add-only, so a correction you make survives a retriage. An edit that changes
nothing material is skipped.
Feature requests and questions stop here. Engineers build features; the bot will not, and the label below is refused on them.
Fix
Apply the referentautomat label, or comment @referentautomat fix (start of
a line, org member or collaborator), on a bug or chore issue. Claude Code
(Opus 5) branches off dev, makes the change, runs Biome, the type check, and
the targeted tests, and commits. A deterministic step then checks the diff
against issue-policy.json (denied paths, at most 15 files and 400 lines) and
only then pushes and opens a draft PR into dev titled
<type>: <subject> (REF-N) with Closes #N and Ref REF-N in the footer.
The model has no git push and no gh. The token in the model's environment
is a read-only App token, and its job holds no other. The branch, OUTCOME.json
and PR_BODY.md leave as an artifact; a separate job checks out a clean dev,
imports the branch, and runs the policy check, push, PR creation, and labels
with the write token. A diff that breaks policy is never pushed; the issue gets
run:manual naming the rule.
Bot PRs are not exempt from the Linear attachment check.
Read the evidence in the PR, mark it ready when it holds up, and merge it as
you would any PR. If it was wrong, apply run:flawed on the issue when you
close or rewrite the PR; that label is the data the deny list is widened from.
Outcomes
| Label | Set by | Meaning |
|---|---|---|
run:success | bot | draft PR opened |
run:blocked | bot | could not proceed for reasons outside the diff; see the comment |
run:manual | bot | policy declined: feature kind, no kind label yet, denied path, over budget |
run:flawed | you | the bot said success but the PR was wrong |
The trigger label comes off after every run, so a retry is a deliberate act.
Commands
| Comment | Effect |
|---|---|
@referentautomat triage | Retriage now. |
@referentautomat fix | Draft a fix (same rules as the label). |
Rationale and the scope-widening rule: ADR 0002.
Configuration
| Secret or variable | Purpose |
|---|---|
REFERENT_APP_ID | The App's ID, for minting installation tokens. |
REFERENT_PRIVATE_KEY | The App's private key. |
LINEAR_BOT_API_KEY | Linear key for candidate search and for the QA checklist issue. The release pipeline's LINEAR_ACCESS_KEY does not authenticate against GraphQL. The checklist needs write access, so the key must be able to create and update issues in the QA team. |
LINEAR_QA_TEAM_KEY | Repository variable. The Linear team key (QA) the checklist issue is created in. The QA Checklist label must exist in that team or the workspace. Unset skips the qa job entirely, so an unconfigured repo shows a skipped check rather than a failure. |
LINEAR_QA_PROJECT_ID | Repository variable, optional. The Linear project the checklist issues are collected in (Release QA). Unset leaves each issue outside any project. Applied on creation only. |
GH_BOT_AI_GATEWAY_API_KEY | AI Gateway key for the notes draft. |
REFERENTA_BOT_DRY_RUN | Repository variable. See below. |
REFERENTA_BOT_LIVE_STAGES | Repository variable. Comma list of issue stages allowed to write (triage, fix). Unset means every stage is in dry run. Independent of REFERENTA_BOT_DRY_RUN. |
The installation needs Contents: read and write (the tag push and the GitHub Release), Pull requests: read and write (bodies, labels, comments), and Issues: read and write (labels are an issues-API concern on GitHub). Every job mints a short-lived installation token at the start of its run; nothing holds a long-lived credential.
The Dry-Run Switch
REFERENTA_BOT_DRY_RUN is the kill switch. An unset variable means dry run,
so a capability has to be switched live deliberately. While it is true the bot
narrates its intended writes into the job log and into its comment, prefixed
dry run, and performs none of them. Informational findings still post in
either mode.
In dry run the notes head is never written, so the automatic refresh on pushes
into the RC is disabled: it would otherwise redraft on every push and discard
the result. The draft still runs, narrated, when the RC is opened or reopened
and on @referentautomat notes.
The same holds for the QA checklist: in dry run the RC body is never stamped,
so the automatic refresh on pushes is disabled and the job narrates its
intended Linear writes on open, reopen, the QA Required label being added,
and @referentautomat qa only.
To pause the bot without touching a workflow file, set the variable to true in
Settings ▸ Secrets and variables ▸ Actions ▸ Variables. It takes effect on
the next run.
Working On The Bot
The scripts are plain Node ESM with no dependencies and no build step, so
node --test runs the suite directly:
node --test .github/referenta-bot/scripts/*.test.mjsreferenta-bot-tests.yml runs the same command in CI, but only when the
.github/referenta-bot/ tree itself changes. The scripts sit outside the pnpm
workspaces and are excluded from pnpm test, so without that path filter the
suite would never execute.
The pure logic lives in its own modules for exactly this reason:
match-labels.mjs (globs and label matching), pr-policy.mjs (title
validation, wrong base, footer stripping), linear-refs.mjs (identifier
extraction and candidate ranking), notes-region.mjs (the marked region and
the notes head), and release-policy.mjs (which event or comment asks for
which capability, and who may ask). They do no I/O, so they are unit-testable
without a GitHub token.
The orchestration scripts (run-pr-checks.mjs, audit-rc.mjs,
apply-notes.mjs) each export their body as a function that takes its I/O
(github, git, linear) as arguments, with a CLI guard at the bottom that
supplies the real ones. testing/fake-fetch.mjs is a fetch double that
routes by method and path and records every call, so a test can drive a whole
script and assert on exactly the writes it made. Keep new logic in that shape:
a pure module, a thin script that exports its body, and a test that runs it
against the double.
Workflow changes are worth linting locally with actionlint, which CI does not run:
actionlint .github/workflows/*.yml.github/actionlint.yaml is committed for this, declaring the
blacksmith-4vcpu-ubuntu-2404 runner label and disabling the config-variable
check. Without it every job in the repo is reported as using an unknown runner
and the real findings are buried.
Troubleshooting
A label was not applied. Check that the label exists in the repo and that
your path is actually covered by a glob in label-map.json. Remember * does
not cross a /.
The notes job did not run. It requires an exact <Prefix>: vX.Y.Z title.
The PR check reports the near-miss; fix the title and the next push into the RC
drafts them, or comment @referentautomat notes to draft them now. If a
command did nothing, check the Release Candidate run for that comment:
it logs who commented and why the command was ignored.
The audit says the notes are behind dev. The notes head names the dev
commit the region was drafted for, and PRs merged after it are listed. The
next push into the RC redrafts automatically; @referentautomat notes
does it now.
The notes came back wrong. Edit the region by hand. Your text wins on the
next refresh. If the draft is systematically off, the fix belongs in SKILL.md
rather than in one release's body.
A finding will not clear. The checks re-run on edited and synchronize.
If a stale comment is left after the underlying issue is resolved, re-run the
Referentautomat PR workflow for that PR from the Actions tab; the sticky
comment is deleted when there is nothing left to report.
The tag or release came out under a person's name. That means the release job did not use an app token, so check whether it was reverted to a personal one. It cannot be a silently-lost permission: the job requests Contents write explicitly at the mint step, so losing the grant fails the job there, naming what is missing, rather than falling back to another identity.