Running core: what recovers after a crash, what keeps node state honest, and what to watch.
Core journals into the store, under /wal/{service address}/{seq} — the same address it registers
itself under, so an instance’s journal and its liveness key always agree. The journal is intent,
not state: an entry is written before the risky step and deleted once the step has reached a
stable outcome. So whatever survives under the prefix after a crash is exactly the work that was
in flight, and DisasterRecover replays that prefix at startup, before the gRPC server starts
serving. seq is an in-process counter, seeded on start from the highest key already there and
zero-padded, so replaying the keys in order replays the entries in order.
| Event | Written before | Replay does |
|---|---|---|
allocate-workload |
resources are allocated on a set of nodes, a workload is removed, or a realloc starts | re-derives each node’s usage from its actual workloads (NodeResource with fix) |
create-workload |
the engine is asked to create a workload, and before one is removed | removes the workload — from the store if it is there, otherwise off the engine, found by name when the entry has no ID yet |
replace-workload |
the old workload of a replace is removed | removes the old workload if the new one reached the store, releasing nothing, because the new one inherited its resources |
realloc-workload |
a realloc mutates plugin usage, metadata and engine limits | re-applies the stored engine params to the workload |
create-processing |
an in-flight deploy counter is written | deletes the stale counter, so it stops inflating deploy counts forever |
create-lambda |
a RunAndWait workload starts |
waits for it to exit, then removes it |
Each replayed handler gets a 32-second deadline. Entries whose handler is unknown are logged and skipped; entries that fail are logged and left in place for the next start.
Because the journal is in the store, any instance can finish another’s work. The active node
status watcher keeps the live service addresses from the store’s service stream and, every
grpc.service_heartbeat_interval, replays the journal of every /wal/ prefix whose address is no
longer registered. It holds /wal-replay/{address} while it does, so two instances that disagree
about who is dead still cannot replay one journal twice; the handlers are idempotent, so a second
pass is harmless anyway.
An instance whose service key merely flapped — a lost lease, not a dead process — can therefore
have its in-flight entries replayed under it. Every handler is written to converge in that case,
but it is the reason the takeover waits for the address to leave the registry rather than probing
it directly. GetNodeResource(fix: true) remains the manual reconciler.
The two NodeResource/PodResource calls are the manual counterpart: they list a node’s
workloads, ask the plugins for capacity and usage, inspect each workload on the engine, and report
diffs. With fix: true, the plugins rewrite usage from the workloads that actually exist.
Every core instance starts a node status watcher, but only one is active at a time. They compete
for an ephemeral key at /selfmon/active with a ha_keepalive_interval TTL; the winner runs, the
losers retry every second and log once a minute. If the winner dies, its key expires and another
takes over.
The active watcher:
/status:node/{nodename} key expired, meaning its agent stopped
reporting — calls SetNode with workloads_down, marking that node’s workloads dead.It deliberately ignores nodes coming back up: that transition is owned by the agent, which
re-reports the node and its workloads. Nodes registered with test: true skip the liveness check
and count as alive unless they are also bypassed.
Separately, the engine cache subscribes to the same stream and evicts every cached engine client belonging to a node that went down — see Engines.
Set profile to a host:port to expose an HTTP server with /metrics in Prometheus format and
the standard net/http/pprof handlers. With profile empty, neither is served.
Scraping /metrics is not passive: the handler first walks every node, sends its up/down gauge
and asks the resource plugins for that node’s metrics, then serves the registry. A scrape
therefore costs one pass over the cluster and is bounded by global_timeout — keep the scrape
interval well above it.
Two metrics are core’s own:
| Metric | Type | Labels | Meaning |
|---|---|---|---|
core_deploy |
counter | hostname |
Workloads this instance has scheduled |
pod_node_up |
gauge | hostname, podname, nodename |
1 when the node is neither down nor bypassed |
Everything else is declared by the resource plugins through GetMetricsDescription, so a cluster
running the cpumem plugin plus gpu and storage plugins exposes their gauges too. When a node is
removed, or turns out to be invalid, its label set is deleted from the collectors so stale series
do not linger.
If statsd is set, every metric is mirrored there as well, over UDP, under keys like
core.<hostname>.deploy.count and pod.node.<nodename>.up. The statsd connection is lazy and
failures are logged, never fatal.
Setting auth.username installs a unary and a stream interceptor on every rpc. The check is
deliberately simple: the request’s gRPC metadata must contain an entry whose key is the
configured username and whose value is the password. There is no TLS on the core listener and
the client does not require transport security, so this is a shared secret on a trusted network,
not a public-internet authentication scheme. Put core behind something else if the network is not
trusted.
With auth.username empty, the API is open to anyone who can reach the port.
Setting sentry_dsn initializes the Sentry client. Anything logged at error or fatal level is
reported with its stack and, where core knows them, the caller’s address and trace ID as tags.
Goroutines spawned through core’s own helper capture panics to Sentry before re-raising them, and
the process flushes on exit. Empty means no Sentry.
The same profile listener serves net/http/pprof:
go tool pprof http://<core>:12346/debug/pprof/profile?seconds=30
go tool pprof http://<core>:12346/debug/pprof/heap
curl http://<core>:12346/debug/pprof/goroutine?debug=2
Core also sets GOTRACEBACK=crash in its systemd unit and raises LimitCORE, so a panic leaves a
core dump behind.
On SIGINT, SIGTERM or SIGQUIT core stops accepting new work, unregisters its
/services/{addr} key so clients stop routing to it, gracefully stops the gRPC server, and then
waits for in-flight streaming tasks to finish before exiting. Long calls — a build, a
RunAndWait, a log stream — hold shutdown open, which is why the systemd unit allows 1200
seconds.
Each instance’s identifier is the SHA-256 of its marshalled config, and every workload it creates
is labelled eru.coreid with that value. Instances sharing a config share an identity by design —
it identifies the cluster configuration a workload belongs to, not the process that created it.
Info returns it.