This document covers what you do when a yuno does not behave correctly. It
explains how to enable the traces that show you what happens, where the output
goes, and how to follow one message through several yunos. It also explains the
part that the centralized log aggregator (logcenter) plays.
This document is the companion to YUNO_LIFECYCLE.md.
That one covers how the agent manages yunos. This one covers how to look
inside them.
1. Mental model¶
Three observation layers are independent. A confusion between them is the first source of frustration:
| Layer | Question it answers | How you turn it on |
|---|---|---|
| Log (severity) | “Did something bad happen?” | Always on. Filter by severity in the log file. |
| Trace (categories) | “What was the system doing a moment ago?” | set-global-trace / set-gclass-trace / set-gobj-trace — off by default. |
| Audit | “What commands did operators run on this yuno?” | Always written when use_audit_command_file=true. |
You configure four destinations per yuno, with daemon_log_handlers in the
yuno config JSON:
┌──────────────────────────┐
severity logs │ file handler │ → <yuno dir>/logs/<mask>.log
+ traces ──────►│ (rotatory, ~8 MB) │
└──────────────────────────┘
┌──────────────────────────┐
│ udp handler │ → udp://host:port
│ (default :1992) │ → typically the logcenter yuno
└──────────────────────────┘
┌──────────────────────────┐
│ stdout (console mode)│ → terminal when not daemonised
└──────────────────────────┘
┌──────────────────────────┐
│ remote_log over │
│ ievent / websocket │ → SPA "dev panel" (live viewer)
└──────────────────────────┘One log line can go to all four destinations at the same time. No destination is “the” log. They are different sinks.
2. Severity levels (gobj_log_*)¶
These are the calls every gclass uses to record events. They are not
traces — they fire regardless of trace settings. The six public ones are
defined in kernel/c/gobj-c/src/glogger.c:
| Function | Priority | At |
|---|---|---|
gobj_log_alert | LOG_ALERT | glogger.c:499 |
gobj_log_critical | LOG_CRIT | glogger.c:514 |
gobj_log_error | LOG_ERR | glogger.c:529 |
gobj_log_warning | LOG_WARNING | glogger.c:544 |
gobj_log_info | LOG_INFO | glogger.c:559 |
gobj_log_debug | LOG_DEBUG | glogger.c:574 |
glogger.c declares two more channels that are not syslog channels:
LOG_AUDIT(8) — the framework writes these lines without the standard header. They record the command audit. A tool that filters on the standard timestamp format skips them, so read the audit file directly.LOG_MONITOR(9) — the monitoring tools use this channel.
Per-yuno HARD RULE (see CLAUDE.md): every
error-return path calls gobj_log_error or carries an
// Error already logged comment. If you cannot find the error in the log,
that yuno has a bug. The log did not lose it.
3. Trace categories¶
A trace is the running commentary that the framework can emit. It is off by default. It is noisy, so enable it only when you need it, and disable it when you finish.
3.1 Global trace levels¶
Defined in s_global_trace_level[16] at kernel/c/gobj-c/src/gobj.c:
| Bit | Name | Emits when |
|---|---|---|
| 0 | machine | Every FSM event dispatch + every state change. The big one. See §6. |
| 1 | create_delete | gobj created / destroyed |
| 2 | create_delete2 | Same as above, plus the kw payload |
| 3 | subscriptions | gobj_subscribe_event / gobj_unsubscribe_event |
| 4 | start_stop | gobj_start / gobj_stop |
| 5 | ev_kw | Dump the kw JSON payload on every event dispatch (huge volume) |
| 6 | authzs | Authorization checks |
| 7 | states | State changes (subset of machine) |
| 8 | gbuffers | gbuffer alloc / free / realloc |
| 9 | timer | One-shot timer fires |
| 10 | fs | Filesystem ops — including timeranger2 appends |
| 11 | liburing | io_uring submit / complete |
| 12 | timer_periodic | Periodic timer fires (separate from timer to avoid spam) |
| 13 | liburing_timer | io_uring-backed timers |
| 14 | commands | gobj_command invocations |
These are global bits. When you enable one, it affects every gobj in the yuno.
3.2 Per-gclass trace levels¶
Each gclass declares its own up-to-16 levels in s_user_trace_level[16].
Example: c_tcp_s.c
enum {
TRACE_LISTEN = 0x0001,
TRACE_NOT_ACCEPTED = 0x0002,
TRACE_ACCEPTED = 0x0004,
TRACE_TLS = 0x0008,
};
PRIVATE const trace_level_t s_user_trace_level[16] = {
{"listen", "Trace listen"},
{"not-accepted", "Trace not accepted connections"},
{"accepted", "Trace accepted connections"},
{"tls", "Trace tls"},
{0, 0},
};The names are gclass-specific. Common ones across runtime gclasses:
C_TCP,C_TCP_S:traffic,connect,tls,listen,accepted,not-acceptedC_PROT_HTTP_SR,C_PROT_HTTP_CL:trafficC_IEVENT_SRV,C_IEVENT_CLI:ievents,ievents2(the second dumps full kw)C_WEBSOCKET: gclass-specificdebugfor HTTP-upgrade handshakes
To see what a gclass offers, run get-gclass-trace gclass=<X> (see §4).
3.3 Per-gobj trace levels¶
These levels are the same as the per-gclass levels, but they are scoped to one
gobj instance. They are useful when you have ten TCP connections and you want
the trace of one connection. API:
gobj_set_gobj_trace() at kernel/c/gobj-c/src/gobj.c:11256.
3.4 The no_trace parallel system¶
For every “set trace” command there is a “set no-trace” counterpart. The
framework subtracts the no-trace mask from the effective trace mask. So
you can enable a noisy level globally, then silence it on specific gclasses or
gobjs.
Functions: gobj_set_global_no_trace() at gobj.c:11711,
gobj_set_gclass_no_trace() at gobj.c:11617, gobj_set_gobj_no_trace() at gobj.c:11746.
3.5 Deep trace mode¶
gobj_set_deep_tracing(level)
enables all traces, and the masks do not apply. There is no
ycommand for it. It is available only in the C API, and
the framework uses it internally for emergency dumps. Do not use it unless you
can accept the volume.
4. Turning traces on and off¶
All commands go to the yuno itself, addressed to its __yuno__ service.
Handlers in kernel/c/root-linux/src/c_yuno.c:
# discover what a gclass offers
ycommand -c 'command-yuno id=<yuno> service=__yuno__ command=get-gclass-trace gclass=C_TCP_S'
# enable / disable
ycommand -c 'command-yuno id=<yuno> service=__yuno__ command=set-global-trace level=machine set=1'
ycommand -c 'command-yuno id=<yuno> service=__yuno__ command=set-gclass-trace gclass=C_TCP_S level=traffic set=1'
ycommand -c 'command-yuno id=<yuno> service=__yuno__ command=set-gobj-trace gobj=<short_name> level=machine set=1'
# silence (no_trace) — per gclass or per gobj
ycommand -c 'command-yuno id=<yuno> service=__yuno__ command=set-gclass-no-trace gclass=C_TIMER level=periodic set=1'
ycommand -c 'command-yuno id=<yuno> service=__yuno__ command=set-gobj-no-trace gobj=<short_name> level=machine set=1'
# inspect current state
ycommand -c 'command-yuno id=<yuno> service=__yuno__ command=get-global-trace'
ycommand -c 'command-yuno id=<yuno> service=__yuno__ command=get-gclass-trace gclass=C_TCP_S'
ycommand -c 'command-yuno id=<yuno> service=__yuno__ command=get-gobj-trace gobj=<short_name>'The short form in CLAUDE.md,
ycommand -c 'command-yuno id=<id> service=__yuno__ command=…', is exactly
this. The shorter form ycommand -c 'set-global-trace …' sends command-yuno
to the yuno that is registered as the default yuno.
set-global-no-tracesilences a global level for every gobj. Each yuno’smain.csets its defaults this way before it creates the yuno, almost always asgobj_set_global_no_trace("timer_periodic", TRUE). That is why a globalmachinetrace does not drown in timer ticks. The command can undo such a default, and the change persists (see Persistence):# see the periodic timer event, and keep seeing it after a restart ycommand -c 'command-yuno id=<yuno> service=__yuno__ command=set-global-no-trace level=timer_periodic set=0' ycommand -c 'command-yuno id=<yuno> service=__yuno__ command=set-global-trace level=timer_periodic set=1'
Persistence¶
Every trace set from the control plane is persisted. None of them are
live-only. set-global-trace, set-gclass-trace and set-gobj-trace all
end in gobj_save_persistent_attrs() on the yuno’s trace_levels attribute,
and their no-trace counterparts on no_trace_levels:
| Command | Saver | Key in the attr |
|---|---|---|
set-global-no-trace | save_global_no_trace | __global_no_trace__ (in no_trace_levels) |
set-global-trace | save_global_trace | __global_trace__ |
set-gclass-trace | save_user_trace | the gclass name |
set-gobj-trace | save_user_trace | the gobj name |
set-gclass-no-trace | save_user_no_trace | the gclass name |
set-gobj-no-trace | save_user_no_trace | the gobj name |
They are re-applied on the next start: set_user_gclass_traces() and
set_user_gclass_no_traces() run from mt_create, set_user_gobj_traces()
right after the children are built.
A saved scope REPLACES what is in force. The global scope, the global
no-trace scope and the scope of each gclass are saved whole, from the
levels in force after the command, and an empty scope is saved as []. At the
next start up, a scope that is in the attr replaces what main() set before it
created the yuno. A scope that was never saved keeps the default of main().
So a default that the user turned off stays off:
"no_trace_levels": {
"__global_no_trace__": [],
"C_TIMER": ["machine"]
}With this attr, timer_periodic is not silenced after a restart, although
main.c silences it, and C_TIMER keeps its machine no-trace. (A scope saved
by an older release is the list of levels that were set one by one; it now
replaces the defaults of main() too.) A gobj-name key, which
reset-all-traces gobj=… writes, keeps its old level-by-level list.
CAUTION: a forgotten set-gclass-trace gclass=C_TCP_S level=traffic set=1
survives a restart exactly like a global one, and it fills your disk. It gives
no message first. Always pair the enable and the disable in the same session.
An entry whose key names a gclass that no longer exists is skipped without a log at start up (
c_yuno.c:4992, a deliberate exception to the no-silent-errors rule: a mistyped gclass name is the common case). So a trace that does not turn on after a restart usually means a typo in the persisted attr. Read it withlist-persistent-attrs, and clear it withremove-persistent-attrs.
4.1 When the yuno never reaches the agent (--global-trace)¶
Every command above travels over the yuno’s control channel to the agent. So
none of them work for the failure that most needs a trace: a yuno that dies,
hangs or fails before that channel is ready. ycommand cannot reach it,
and list-yunos reports running=false even while the process is alive.
Since 7.8.2 the levels can be armed on the command line instead:
# one level, or several — repeatable and comma-separated
auth_bff --config-file='[...]' --global-trace=machine
auth_bff --config-file='[...]' --global-trace=machine,create_delete,start_stop
# what levels exist
auth_bff --global-trace=listThe framework applies them after it registers every gclass, and before the
first service starts, so they cover start up itself. It applies them again
right after it creates the yuno, because the yuno replaces the global scope with
the persisted one: what you ask on the command line wins. A later trace command
saves the global levels in force, and those include the ones from the
command line. An unknown level stops the
yuno with a message that points at list. The yuno does not ignore it.
To reproduce a yuno that the agent launches, take its command line from
running-bin id=<id> or running-keys id=<id>. You can also use the script
that the agent writes at
/yuneta/realms/<realm>/<yuno>/bin/<role>^<id>.sh. Then append the flag.
Two older methods, and their limits:
kill -10 <pid>(SIGUSR1) cycles the global mask0→0x00FF0000→0x0FFF0000→0xFFFF0000→0. It is useful on a process that already runs. It does nothing for a problem during start up.--verbose-log=Nis not a trace switch. It only overrides the stdout log handler’s field bitmask (which fields each line prints).--verbose-log=3prints fewer fields than the config default of 255, which is why it reads as “it does nothing”. Use--global-trace.
4.2 Two warnings that arrive without being asked for¶
Some failures below the framework cannot wait for a trace, so they report themselves:
getaddrinfo() BLOCKED the event loop(gobj_log_warning, msgsetOS, fromyev_loop.c) — name resolution is synchronous and runs inside the loop. A slow resolver therefore stops every gobj in the process, not only the socket that you open. The framework emits the warning with the host and the elapsedmsecwhen the time is more than 1 s. If you see it, the yuno is not slow. It is stopped.YUNETAS static_resolv: …in syslog (journalctl) — theCONFIG_FULLY_STATICresolver writes it. That resolver is below the gobj log and cannot reach it. Three conditions emit the message: anameserverin/etc/resolv.confthat does not answer (rate limited to one message per nameserver per 5 min), a resolution of more than 1 s, or a failed allocation.
A dead first nameserver costs ~6 s (A + AAAA timeouts) on every lookup,
so a yuno that opens many channels can spend minutes in start up. The resolver
caches answers since 7.8.2, which limits the cost to the first lookup. But the
correction belongs in the node’s /etc/resolv.conf.
5. Reading the logs¶
5.1 File paths¶
Per-yuno log file, built by yuneta_log_file() at
kernel/c/root-linux/src/yunetas_environment.c:
<work_dir>/<domain_dir>/logs/<filename_mask>. work_dir is /yuneta. The
agent gives each yuno that it runs its own domain_dir
(build_yuno_private_domain() in c_agent.c):
/yuneta/realms/<realm_owner>/<realm_name>.<realm_role>.<realm_env>/<role>^<id>/logs/<filename_mask>For example, /yuneta/realms/artgins/artgins.yunetacontrol.com/controlcenter^1996/logs/controlcenter-4.log.
The agent itself and the utilities run by hand have a fixed domain_dir in
their main.c: /yuneta/realms/agent/agent/logs/ for yuneta_agent,
/yuneta/realms/agent/<utility>/logs/ for ycommand, ybatch and the
others. There is no /yuneta/logs/ directory.
The mask is the value that you set in
daemon_log_handlers.<handler>.filename_mask (see §5.4). By convention it is
<role>-W.log, where the day of the week (1 = Sunday … 7) replaces the
W: on a Wednesday the yuno writes <role>-4.log.
Active log discovery:
ls -lt /yuneta/realms/*/*/*^<id>/logs/
tail -f /yuneta/realms/*/*/*^<id>/logs/<latest>.log | grep -a "keyword"5.2 Log line format¶
Every log record is a JSON object built in glogger.c. Fields
added automatically by discover() at glogger.c:1326:
| Field | Source |
|---|---|
timestamp | current_timestamp() |
priority | LOG_ERR / LOG_WARNING / … |
node_uuid | host node identity |
process | yuno binary name |
hostname | from gethostname |
pid | process id |
gclass | the gclass that emitted the line |
gobj_name | the gobj instance name |
state | current FSM state of that gobj |
gobj_full_name | dotted path (only if gobj_full_name trace is on) |
id | sequence id |
msgset, msg | the "msgset","msg" pair every gobj_log_* call passes |
any key,value | extra fields the caller passed |
Searching is JSON-friendly:
grep -a '"priority":3' <yuno dir>/logs/<file>.log # all errors
grep -a '"gclass":"C_TCP_S"' <yuno dir>/logs/<file>.log # one gclass
grep -a '"msg":"Event NOT DEFINED in state"' … # the canonical FSM bug5.3 Rotation¶
The rotatory library makes the file name from the mask and
the local date, so the file changes at midnight (the first record after it). With the W mask,
the first record of a new day, and the first record after the yuno starts,
empties the file of the same week day when it was last written before today:
that is last week’s file. So a yuno keeps 7 days of log, and the mask is
the retention. A file written today is appended to, so a restart keeps it, and
a clock set back across midnight empties nothing. A mask with no date letter
(a fixed name such as logcenter.log) is never emptied by a date. The table of
which masks empty a file is in File names and rotation.
The library also rotates the file when it becomes larger than a size threshold
(default 8 MB, counted in bytes, configurable via max_megas_rotatoryfile_size,
entry_point.c):
it renames the file to <name>.OLD (a previous .OLD is removed) and starts
the file again. There is no cron. Both rotations happen on the next write.
The file is checked once for each record, never between the pieces of a record,
so a record is never split between two files. A log file removed by hand is
created again by the next record. At a new day the size of the file of the day
before is not checked: that file is left as it is, even over the limit, and its
.OLD stays. (In 7.25.4, when the last piece of a day took its file over the
limit, the first record of the next day renamed that file to .OLD and removed
the .OLD of that day.)
A log file RENAMED by another program (a logrotate with its default
create mode) is not noticed: the yuno goes on writing into the renamed file
until its next new file (up to 7.25.4 it was noticed). To rotate a yuno log
from outside, copy and truncate it (logrotate with copytruncate), or
remove it:
cp /yuneta/realms/agent/agent/logs/yuneta_agent-4.log /tmp/ \
&& truncate -s 0 /yuneta/realms/agent/agent/logs/yuneta_agent-4.logA full disk stops only that file. Every 100 records the handle checks the
free space of its own disk; below min_free_disk_percentage (default 10%) it
drops its records, and it writes again by itself when the space is back. It
prints one line to stdout and syslog when it stops and one when it resumes
(up to 7.25.4 one full disk stopped every file log of the process until it
was restarted):
rotatory(): stop logging to '/yuneta/realms/agent/agent/logs/yuneta_agent-4.log' because full disk: 9% free (<10%)
rotatory(): logging to '/yuneta/realms/agent/agent/logs/yuneta_agent-4.log' again: 12% free (>=10%), 5210 records were droppedSee a full disk in the rotatory page.
With the defaults, one yuno uses at most 7 × 2 × 8 MB = 112 MB of log (each piece ends with the record that took it over 8 MB). Up to 7.25.4 the size was counted in whole megabytes, so a piece rotated only at 9 MB: 126 MB.
5.4 Where to configure handlers¶
In the yuno’s config JSON, under environment.daemon_log_handlers (or
console_log_handlers in non-daemon mode), parsed at
kernel/c/root-linux/src/entry_point.c:
"environment": {
"daemon_log_handlers": {
"to_file": {
"handler_type": "file",
"filename_mask": "mqtt_broker-W.log",
"handler_options": 255
},
"to_udp": {
"handler_type": "udp",
"url": "udp://127.0.0.1:1992",
"handler_options": 255
}
}
}handler_options is a bitmask of LOG_HND_OPT_* (glogger.h) that selects
which severities the handler accepts. 255 accepts all of them. If you clear
bits, the handler drops DEBUG, INFO, AUDIT and the other severities.
To add or remove handlers at run time, use the add-log-handler and
del-log-handler commands of c_yuno.c.
5.5 The agent’s audit files¶
yuneta_agent writes every command that it runs to an audit file, one JSON
record for each command (use_audit_command_file, on by default). The console
writes (write-tty, one command for each keystroke) are the exception: they
make one record for each burst (see Console writes). The files
are in the audit/ directory of the agent realm:
/yuneta/realms/agent/agent/audit/266-23_09_2026.log # ZZZ-DD_MM_CCYY.log
/yuneta/realms/agent/agent/audit/266-23_09_2026.log.OLD.1 # the first 500 MB of a big day
/yuneta/realms/agent/agent/audit/266-23_09_2026.log.OLD.2 # the next 500 MBWhat a record holds¶
The record is built by audit_record_build()
(yunos/c/yuno_agent/src/audit_record.c). There are two forms, and a third
one for the console writes.
A read-only command gets only the command, the date and the user:
{"command":"list-yunos","date":"2026-09-24T10:00:00.000000000+0200","user":"claudia@artgins.com"}The command tables do not say which commands are read-only, so the list is by
name: the prefixes list-, view-, get-, info-, dir-, and help,
authzs, ping, node-uuid, top, top-services, services, stats,
stats-agent, stats-yuno, authzs-yuno, treedbs, treedb-info, topics,
desc, descs, system-schema, schema-file, saved-schema,
diff-schema, jtree, nodes, node, instances, hooks, links,
parents, children, pkey2s, snaps, snap-content, print-role,
print-tranger, check-json, check-realm, cert-expiry-status,
cert-sync-status, global-variables, running-keys, running-bin,
users, accesses, roles, user-roles, user-authzs.
A command that has a __reset__ value (in the command text or in the kw) is
not read-only: stats=__reset__ sets the counters of a yuno to zero. So
stats-yuno id=gate_mqtts stats=__reset__ (the reset button of gui_agent) gets
the full record, with its source, also when the __reset__ comes in the kw.
command-yuno and command-agent are judged by the command that they carry,
and the record names it: "command":"command-yuno command=view-attrs". The
carried command is taken from the same place as the handler takes it: the last
command= of the command text, else command of the kw. So a kw
command=list-yunos with a text command='delete-node …' runs delete-node
and is recorded as delete-node, with the full record.
The key must be exactly command. The command parser matches a key of the text
in any case, but it stores the value under the key as it was typed, and the
handler reads command only. So COMMAND=list-yunos is not the command that
runs: with a kw {"command": "delete-yuno id=gate"}, the command
command-yuno id=gate COMMAND=list-yunosruns delete-yuno, and it is recorded with the full record (the kw, the
source). A key command in another case, in the text or in the kw, always
gives the full record: it is never taken as read-only. The same rule gives the
console of a write-tty: the handler reads name, so NAME=decoy does not
change the console of the record.
These commands keep the full record: read-file, read-json and
read-binary-file (they read files of the node), check-user-pwd, and anything
that opens something (open-list, open-treedb, …).
Every other command gets the command, the date, the user, the source and the parameters:
{"command":"install-binary id=auth_bff content64='<33554432 bytes sha256:401b36b9e4f91e967e815f96fd293cb384222d3e73e1b9fcf759d5ecfadfdbf8>'",
"date":"2026-09-24T10:00:00.000000000+0200",
"user":"yuneta",
"source":{"hops":[{"role":"ycommand","yuno":"","service":"ycommand",
"user":"yuneta","host":"dev-laptop"}]},
"kw":{"__username__":"yuneta"}}The same command sent from gui_agent through the controlcenter has two hops, nearest first:
"source":{"console_purpose":"statnodes",
"hops":[{"role":"controlcenter","yuno":"artgins.com","service":"top-16",
"user":"yuneta_agent@artgins.com","host":"artgins"},
{"role":"gui_agent","yuno":"gui_agent_yuno","service":"agent_link",
"user":"claudia@artgins.com","host":"544f1345-65a6-455c-aa94-62b6b020b5c5"}]}useris the end user (__username__of the kw), or the user of the nearest hop.sourcereplaces__md_iev__. It keeps the console purpose (if any) and, for each inter-yuno hop (nearest first), the role, the yuno and the service of the sender, its user and its host. The address of the peer is not in the kw and is not written. A command typed on the node (localycommand) has one hop, with the roleycommand, the user of the shell and the host name of the node. Only a command that the agent sends to itself has nosource. A field of a hop that is not a string is written as"", and the agent logs one WARNING (msgsetProtocol, no stack): it is data of a peer. For example a hop that arrives as{"src_role": 7}is written"role": ""with the log line “Audit: bad md_iev from a peer, written as empty”.user, each field of a hop,console_purposeand the console name of a console write are data of a peer too: each one is redacted like any other string (a secret, aBearertoken, a JWT in it becomes<redacted>), and one longer than 1024 bytes is written as<N bytes, not scanned, sha256:HEX>. A hop whoseuseris a JWT is recorded as{"role":"gui_agent","yuno":"gui_agent_yuno","service":"agent_link","user":"<redacted>","host":"…"}Up to 7.25.4 the record was the whole kw as it came,
__md_iev__included, and nothing in it was redacted.A
content64is never written. Everywhere (in the command text, whereycommandputs it, with or without blanks around the=, and in any kw key namedcontent64), the value is replaced by<N bytes sha256:HEX>: the size and the sha256 of the DECODED content. The sha256 is the one of the binary, so you can check it on the node:sha256sum /yuneta/repos/yuneta/utils/auth_bff/*/auth_bff # /yuneta/repos/<domain>/<class>/<role>/<version>/<role>A value that is not base64 (a path, for example) is replaced by
<N chars, not base64, sha256 of the text:HEX>. A bad base64 gives no error in the audit: the command itself answers the error.A secret is never written. Its value is replaced by
<redacted>. A secret is a parameter whose name holds, in any case:passw,pwd,passphrase,secret,token,jwt,bearer,authorization,cookie,credentialorsalt,or, with
_,-,.and blanks taken out,apikey,sessionid,sessionkeyorauthdata,or both
privandkey.
For example
password,user_passw,client_secret,kc_admin_client_secret,access_token,api_key,x-api-key,http_cookie,__session_id__,auth_data,visitor_salt,private_key. The list comes from a scan of every attribute and command parameter of the SDK and of the projects. The path of a key or a certificate (ssl_certificate_key), a certificate (cert_pem) and the ids of treedb (pkey,rkey) are not secrets, and stay.cookie_domainis redacted too: it holdscookie.Also the
valueof awrite-attrwhoseattributehas such a name, also whenattributeandvaluecome in a JSON object of the kw. Here the keys are read in any case (ATTRIBUTE=api_key VALUE=…is redacted too). That is stricter than the handler, which reads onlyattributeandvalue: the parser gives it{"ATTRIBUTE": …, "VALUE": …}, keys as typed. A redaction can be stricter than the parser, never looser. The same holds for a JSON text of a write-attr given in a string of the kw, and the names are read with their JSON escapes decoded, at any level: both of these are recorded with"value":"<redacted>":update-node topic_name=x cfg='{"attribute":"api\u005fkey","value":"hunter2"}' update-node topic_name=x cfg='{"attr\u0069bute":"api_key","value":"hunter2"}'A secret value that ends where a
"…"value ends is only the secret: the closing quote and what follows it stay. This commandcommand-yuno id=x command="write-attr attribute=api_key value=hunter2" n=1is recorded as
command-yuno id=x command="write-attr attribute=api_key value=<redacted>" n=1Also the token after
Bearer, the credentials afterBasicwhen they are base64 ofuser:password, and anything with the shape of a JWT (eyJ…, three parts joined by.; a.after it, as at the end of a sentence, is not a fourth part), wherever they are. This applies to a kw key at any depth, toname=valuein any string (quoted or not, with blanks around the=), to the command carried bycommand-yuno, and to"name": valuein a JSON given as text. The name of a JSON key is read with its escapes:"pass\u0077ord"ispassword.The value of a secret
name=valueis its whole word, as the shell reads it: quoted pieces are taken with their blanks,\"inside"…"does not end the value, and the shell’s'\''does not either. A JSON string value of a secret goes on to its own closing quote, also when it holds a'inside a'…'value. So these commands:set-user-pwd username=bob password="a\" hunter2" n=1 set-user-pwd username=bob password='it'\''s hunter2' n=1 update-node topic_name=x cfg='{"password":"it's hunter2"}' n=1are recorded as
set-user-pwd username=bob password="<redacted>" n=1 set-user-pwd username=bob password='<redacted>' n=1 update-node topic_name=x cfg='{"password":"<redacted>"}' n=1The audit takes the value whole even when the parser later refuses the command: the record is written before the parser runs.
The word
Basicalone, or followed by base64 that is notuser:password, is not a secret and stays:update-node topic_name=x data='Authorization: Basic dXNlcjpwYXNz' # dXNlcjpwYXNz: base64 of user:passis recorded as
update-node topic_name=x data='Authorization: Basic <redacted>'A JSON text given inside a JSON text (a treedb column that holds JSON text, for example) has its quotes escaped, and it is redacted too, at any depth up to 8 levels: a quoted run that holds a backslash is scanned again with its escapes decoded. The redacted run is written back with its escapes, so the record still holds the same JSON text:
update-node topic_name=x content='{"cfg":"{\"password\":\"hunter2\",\"n\":1}"}'is recorded as
update-node topic_name=x content='{"cfg":"{\"password\":\"<redacted>\",\"n\":1}"}'This does not depend on the quotes that come before the JSON text. Every quote that is not escaped begins a quoted run, so a double-quoted parameter or a stray quote earlier in the command does not hide it. This command is recorded with
\"password\":\"<redacted>\"too:update-node topic_name=x id="a b" content='{"cfg":"{\"password\":\"hunter2\"}"}'The audit judges what it sees as text, whatever the parser does with it later. When no quoted run holds the escaped JSON whole, the key is read by its shape:
\"name\":, with the backslashes of its level. The key is the whole string back to the quote of that level, so\"secret key\":and\"password/db\":are secrets. Its value is then taken with the quotes of that level. Two examples: the parser endsx="{\"password\":\"hunter2\"}"at the first\", and'{\"password\":\"hunter2\"}'has no quote of its own. Both are recorded with\"password\":\"<redacted>\". For the same reasonpassword=\"two words\"is recorded aspassword=\"<redacted>\".Deeper than 8 levels, a quoted run with a backslash is not written, only its size and sha256 (
<N bytes, not scanned, sha256:HEX>). Up to 7.25.4check-user-pwdandset-user-pwdwrote the password in clear text. For example:{"command":"set-user-pwd username=bob password=<redacted>","date":"…","user":"yuneta", "source":{…},"kw":{"__username__":"yuneta"}} {"command":"check-user-pwd","date":"…","user":"yuneta", "kw":{"username":"bob","password":"<redacted>"}}__command__is left out when it repeats the command text.The command word is taken as the command parser takes it: the first word, without its quotes, looked up in the command table of the agent (any case, and the aliases). So
WRITE-TTY,'write-tty',"Write-Tty"andEV_WRITE_TTYarewrite-tty,1islist-yunos(read-only), andCLOSE-CONSOLEends the bursts of the console writes. The command carried bycommand-agentis looked up the same way. The command carried bycommand-yunoruns in another yuno, whose table the agent does not have: it is compared in lower case.A text that is hard to scan costs no more than its size. The scan of a string is one pass, in linear time: the audit runs before the parser and the authz, on the text that any peer sends. The only recursion is one level for each level of JSON text inside JSON text (at most 8), and each level counts against the same budget. One record scans at most 128 MB: a string beyond it is not scanned and not written, only its size and its sha256 (and, for the command text, its first word before that):
{"command":"set-user-pwd <136314912 bytes, not scanned, sha256:7f3a…>","date":"…","user":"…","kw":{}}
Up to 7.25.4 the record was the command, the date and the WHOLE kw. The sizes, measured with the same commands:
| Command | Up to 7.25.4 | Now |
|---|---|---|
install-binary of a 32 MB yuno (content64 in the command text, as ycommand sends it) | 134,218,673 bytes (the base64 three times) | 541 bytes |
run-yuno through the controlcenter | 891 bytes | 489 bytes |
list-yunos through the controlcenter | 840 bytes | 97 bytes |
On wattyzer, a deploy day wrote 0.6–1.2 GB of audit (5 to 9 binaries) and a normal day 2 KB to 2.6 MB. With this format a deploy day writes a few KB.
Console writes¶
ycommand, ycli and gui_agent send one write-tty command for each
keystroke typed into an agent console, with the keystroke in content64. The
audit keeps only the fact: who, when, which console, how many writes and
bytes. Nothing of what was typed is written, not even a hash: the sha256
of one byte can be read back with a table of 256 entries, so a hash would
give the typed text (and a typed password) back.
The writes of one user (the same user and the same source) into one console
make a burst. A burst lasts 60 seconds from its first write.
The first write of a burst is written at once, alone, and flushed to the file (every audit record is flushed: one
write()for each command, before the command runs). So a crash of the agent cannot lose who typed into a console, even if what was typed made the agent stop.The other writes of the burst make one more record when the burst ends:
dateis the time of its first write,untilthe time of its last one.
{"command":"write-tty","date":"2026-09-24T10:00:00.1+0200","user":"claudia@artgins.com",
"console":"console-1","writes":1,"bytes":1,"source":{…}}
{"command":"write-tty","date":"2026-09-24T10:00:00.4+0200","user":"claudia@artgins.com",
"console":"console-1","writes":212,"bytes":230,"until":"2026-09-24T10:00:58.9+0200","source":{…}}The records of a burst do not overlap, so the sum of their writes and bytes
is the whole burst. A burst ends:
when its 60 seconds have passed, at the next command of any kind (a burst of a console that nobody uses any more waits in memory until then; its first record is already on disk),
when another user or another source writes into its console (so the writes of each user are always recorded apart),
at any
close-console,when the agent stops.
Up to 7.25.4 each keystroke wrote the whole kw, with the keystroke in base64
(about 920 bytes each). Now 1000 keystrokes in one burst write two records,
about 950 bytes. A write-tty carried by command-agent gets its full record, with content64 as
<N bytes> only.
Rotation and retention¶
The mask ZZZ-DD_MM_CCYY.log makes a new name every day. When the file of the
day crosses max_megas_audit_file, it is renamed to the first free
.OLD.<n> (.OLD.1, .OLD.2, …) and a new file begins. No piece of a day
is removed at a rotation: the audit rotatory uses
rotatory_keep_all_old_files(). Up to 7.25.4
each rotation removed the previous .OLD, so a day that crossed the limit
twice lost its first part (on wattyzer, the mornings of 22 and 23 September
2026).
If the directory refuses the rename (chattr +a on it, a read-only bind, a MAC
denial), nothing is removed: the file of the day is kept and grows over the
limit. The agent prints one line (stdout and syslog) and tries the rename again
after 60 seconds of real time (the monotonic clock: a clock set back or forward
does not move the retry), or at the next day, not at every command. It prints
one more line when a rename works again:
_rotatory(): Cannot rename '/yuneta/realms/agent/agent/audit/267-24_09_2026.log' to '/yuneta/realms/agent/agent/audit/267-24_09_2026.log.OLD.1', Permission denied, the file is kept and grows, the rename is tried again every minute
_rotatory(): the size rotation of '/yuneta/realms/agent/agent/audit/267-24_09_2026.log' works againThe attribute audit_keep_days (default 7) is the retention. The agent
removes the audit files older than that number of days:
when it starts, and
when a new audit file begins (a new day, or the size limit). If the new file of a day cannot be opened (no descriptors, a quota, a directory that refuses writes for a moment), the retention runs at the next open that works, once.
The agent opens its audit with exit_on_fail (the last parameter of
rotatory_open()), and each yuno opens its file log the
same way: the process does not start if that file cannot be opened. This
applies to that first open only. A failure after the start never stops the
agent or the yuno. This applies to a new day, a file removed from the
directory, the open after a failed write, and a truncate. The agent prints
one line (stdout and syslog), and the command runs without its audit record.
The next command tries the open again, with no line while it fails. It
prints one line when the open works again, and the retention runs then (see
exit_on_fail is for the open):
_rotatory(): Cannot create '/yuneta/realms/agent/agent/audit/269-26_09_2026.log' file, Too many open files
_rotatory(): '/yuneta/realms/agent/agent/audit/269-26_09_2026.log' is open againUp to 7.25.4 the first failure after the start exited the agent, and the retention of that day did not run.
The second sweep runs inside the write of the first record of the new file,
before that record is written: so it is on the write path of that one command,
once a day (or once for each size rotation). The same file opened again (after
a write that failed, or when the file was removed) is not a new file, and it
runs no sweep. It reads the directory once. It
removes only regular
files with the name shape of the mask (and their .OLD / .OLD.<n>), never a
file of the current day, never a symbolic link, never another file in the
directory (rotatory_remove_old_files()). Each
sweep that removes something writes one INFO line to the agent log:
{"msg": "Old audit files removed", "audit_keep_days": 7, "removed": 3, "megas": 961,
"current_file": "/yuneta/realms/agent/agent/audit/266-23_09_2026.log",
"files": ["258-15_09_2026.log", "258-15_09_2026.log.OLD.1", "258-15_09_2026.log.OLD.2"]}| Attribute | Default | Meaning |
|---|---|---|
use_audit_command_file | 1 | Write the audit files. |
max_megas_audit_file | 500 | Size of one piece of a day, in MB. A bigger day continues in .OLD.<n> pieces. |
audit_keep_days | 7 | Days of audit files kept. 0 keeps all (the behaviour up to 7.25.4). |
min_free_disk_percentage | 20 | Stop writing the audit when the disk has less free space (checked every 100 records), and write it again when the space is back. A new day still runs the retention while the disk is full: the retention is what frees the space. If the new file of that day cannot be opened, the open is tried again every 100 records until it works, and then the retention runs. |
With the new record format the directory is a few MB a week. The retention still bounds it by days, whatever a day writes.
To keep 30 days, set the attribute in the agent config
(/yuneta/agent/yuneta_agent.json) and restart the agent:
{
"global": {
"agent.audit_keep_days": 30
}
}To change it on a running agent (it applies at the next new audit file, and is lost at restart):
ycommand -S __yuno__ -c 'write-attr gobj=agent attribute=audit_keep_days value=30'
ycommand -S __yuno__ -c 'view-attrs gobj=agent attribute=audit_keep_days'An existing large audit/ directory (up to 7.25.4 nothing was removed,
and 19 GB and 90 GB were seen) needs no manual action. The first start of an
agent with this version removes every audit file older than 7 days, logs one
INFO line with the list, and keeps the last 7 days and today. On a slow disk
this first sweep can take some seconds, once. To keep more, set
audit_keep_days before that start.
A clock set back across midnight empties no audit file: the file of the
day before is opened again and the records are appended to it. Up to
7.25.4 it was opened with "w" and emptied, and the file of today was
emptied too when the clock went forward again (see
the rotatory).
A tool that reads the audit must accept both formats: the files written
before the upgrade have the whole kw (with __md_iev__ and the base64) and
no user or source. It must also accept the write-tty records, which have
console, writes, bytes and until and no kw, and <redacted> in place
of a secret.
6. The FSM trace (machine)¶
This is the most useful trace for the behavior of a gobj. It is defined in
glogger.c (trace_machine). Called from the event dispatcher in
gobj.c:
Before dispatch: a
🔜line per event entry.“Event NOT DEFINED” error: a
📛line. This is the canonical failure in which the parent FSM does not declare the event of the child (see CLAUDE.md “CHILD vs SERVICE” section).
After dispatch: a
🔄line per executed event.State change: a
🔀🔀line.
Two output formats, switched by the integer variable trace_machine_format.
Format 1 — one line per transition. THE DEFAULT, on a node and in the
browser alike (gobj.c: trace_machine_format = 1; // 0 legacy, 1 simpler;
gobj-js followed at 7.13.7):
🔜 EV_RX_DATA !!c_tcp :open
🔄 EV_RX_DATA !!c_tcp :open from !!service_main
🔝🔝 EV_ON_MESSAGE c_prot_tcp4h^output-0 :wait_payload
🔝🔄 EV_ON_MESSAGE (EV_ON_MESSAGE) c_channel^output-0Format 1 writes no return line and no state line: the transition line already carries the state it ran in.
Format 0 — legacy, three lines for one transition:
🔜 mach(!!c_tcp), st: :open, ev: EV_RX_DATA, from(!!service_main)
🔄 mach(!!c_tcp), st: :open, ev: EV_RX_DATA, from(c_tcp_s^server)
🔀🔀 mach(!!c_tcp), new st(:closed), prev st(:open)
<- mach(!!c_tcp), st: :closed, ev: EV_RX_DATA, ret: 0Every line of either format is indented by its nesting depth, two spaces per level, so an event sent from inside another one’s action sits under it. That indentation is the only thing that says a transition happened during another — keep it when you render the line anywhere else.
!! before a name means that the gobj is not running at that moment. Two
of them in a row are usually the bug.
Scoping the machine trace¶
The machine trace of a whole yuno gives too much output on anything larger than a toy test. You can make it narrow in two ways:
# only one gclass
ycommand -c 'command-yuno id=<yuno> service=__yuno__ command=set-gclass-trace gclass=C_TCP_S level=machine set=1'
# only one instance
ycommand -c 'command-yuno id=<yuno> service=__yuno__ command=set-gobj-trace gobj=<short_name> level=machine set=1'Both are persisted and are re-applied on the next start, like every other trace set from the control plane (see Persistence). Disable them in the same session.
The same trace in the browser¶
This chapter is about a yuno on a node. But a browser SPA runs the same
kernel, ported to JavaScript. Since @yuneta/gobj-js 7.9.5 the JS
runtime has this level model, with the same names and the same bits. A habit
that you learn here therefore transfers, and you can read two traces side by
side.
gobj_set_global_trace("machine", true); // the big one, same as above
gobj_set_gclass_trace("C_MY_VIEW", "machine", true);
gobj_set_gobj_no_trace(noisy_src, "machine", true); // veto, by the SOURCE
set_log_callback((level, msg) => { ... }); // the trace arrives as `debug`There is no ycommand on that side. The switch is the call above, and the
output goes to the browser console. It goes to any other destination that
set_log_callback() selects, and that is how the dev panel of gobj-ui shows
the machine inside the app.
doc.yuneta.io/navigation runs three demos
with the panel connected to that callback. Read them if you want to see the
lines before you write your own code.
7. Following a message end-to-end¶
Canonical request flow on a typical Yuneta service:
external client
│
▼
┌─────────────┐ gclass trace 'traffic'
│ C_TCP_S │ gobj_trace_dump_gbuf(gobj, gbuf, …)
└──────┬──────┘
│
▼
┌─────────────────┐ gclass trace 'traffic'
│ C_PROT_HTTP_SR │
└────────┬────────┘
│
▼
┌─────────────────┐ gclass trace 'ievents' / 'ievents2'
│ C_IEVENT_SRV │ trace_inter_event2(gobj, prefix, event, kw)
└────────┬────────┘
│
▼
┌─────────────────┐ global trace 'machine' lights up the FSM dispatch
│ service gclass │ gclass-specific traces fire its custom emit points
└────────┬────────┘
│
▼
┌─────────────────┐ global trace 'fs'
│ timeranger2 │ record append + rowid emitted
│ (treedb) │
└────────┬────────┘
│
▼ outbound publish
┌─────────────────┐ gclass trace 'ievents' / 'ievents2'
│ C_IEVENT_SRV │
└────────┬────────┘
│
▼
┌─────────────────┐ gclass-specific trace
│ C_WEBSOCKET │ gobj_trace_dump frames
└────────┬────────┘
│
▼
SPA browserThe correlation id¶
Inter-event messages between yunos carry a metadata block named __md_iev__
inside the kw. Inside it is the ievent_gate_stack — a LIFO of
hops, each entry: {src_yuno, src_service, dst_yuno, dst_service, user, host, …}.
Constant
IEVENT_STACK_ID = "ievent_gate_stack"atkernel/c/root-linux/src/msg_ievent.h.Pushed on outgoing request, popped + reversed on incoming response, at
kernel/c/root-linux/src/msg_ievent.c.
To grep the same transaction across multiple yunos’ logs:
grep -a 'ievent_gate_stack' /yuneta/realms/*/*/*/logs/*.log | grep '<the user or src_yuno you care about>'The framework propagates no automatic UUID for calls that are not ievents. A direct C function call has nothing to grep. The correlation is available only when the message crosses an ievent boundary.
Practical sequence to follow one HTTP request¶
YUNO=my_service_01
# 1. ingress + protocol
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-gclass-trace gclass=C_TCP_S level=traffic set=1"
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-gclass-trace gclass=C_PROT_HTTP_SR level=traffic set=1"
# 2. internal FSM
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-global-trace level=machine set=1"
# 3. broker/topic write
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-global-trace level=fs set=1"
# 4. egress to SPA
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-gclass-trace gclass=C_IEVENT_SRV level=ievents set=1"
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-gclass-trace gclass=C_IEVENT_SRV level=ievents2 set=1"
# trigger the request, capture the noise
tail -F /yuneta/realms/*/*/*^$YUNO/logs/*.log > /tmp/$YUNO.trace &
# … reproduce …
kill %1
# disable everything
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-gclass-trace gclass=C_TCP_S level=traffic set=0"
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-gclass-trace gclass=C_PROT_HTTP_SR level=traffic set=0"
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-global-trace level=machine set=0"
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-global-trace level=fs set=0"
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-gclass-trace gclass=C_IEVENT_SRV level=ievents set=0"
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-gclass-trace gclass=C_IEVENT_SRV level=ievents2 set=0"8. The SPA-side “dev panel” viewer¶
A SPA built on the JS gobj framework can connect to a yuno. Then it can display live what crosses the websocket: the same lines that you see in the local log file, plus the bodies of the ievent messages.
Wire-up¶
C side: nothing special — inter-events flow as usual through
C_IEVENT_SRV→C_WEBSOCKET.JS side: the traffic trace is the gclass level
ievents(orievents2) ofC_IEVENT_CLI(kernel/js/gobj-js/src/c_ievent_cli.js), the same level name asC_IEVENT_SRVon a node. With it on, each message goes totrace_ievent_callback(prefix, iev_msg, direction, size), a yuno attr, or to the console when the attr is empty.The Developer window of gobj-ui (
kernel/js/gobj-ui/src/yui_dev.js) installs that callback inapply_dev_traces(), andinfo_traffic()appends each message into#developer-traffic-logger.
The trace buttons are the yuno’s commands¶
Each button of the window turns trace bits on or off with the trace commands of
the JS C_YUNO, which has the commands and the attrs of the C one
(set-global-trace, set-global-no-trace, set-gclass-trace,
set-gclass-no-trace; trace_levels, no_trace_levels). The yuno persists
them through the functions that the app gives to gobj_start_up(), and
restores them in its mt_create with the same rule as C: a saved scope
replaces the defaults of main.js.
| Button | What it sets |
|---|---|
| Automata | global machine; a second click adds ev_kw |
| Creation | global create_delete |
| Start / Stop | global start_stop |
| Subscriptions | global subscriptions |
| Traffic | C_IEVENT_CLI level ievents |
| Periodic | global timer_periodic, and it clears the global no-trace of the same level |
| I18n | not a trace level: i18next’s debug switch, kept in the browser |
The window never reads a message to decide whether to show it. The state of a button is read from the runtime, so it cannot disagree with what is traced.
Filtering on the SPA side¶
The window filters only what it shows: the direction (in / out / error) and a free text. To change what is traced, use the buttons, which change what the gobjs emit.
Teardown order — the recursion gotcha¶
When the websocket closes, ac_on_close (c_ievent_cli.js:897) fires
EV_ON_CLOSE. set_remote_log_functions redirects the JS log_error and
log_warning calls to the DOM callback. If it is still installed, the callback
captures the warning that the teardown path emits. The callback changes the
DOM, the change can fire more events, and those events log again. The result is
an infinite recursion.
The correction at c_ievent_cli.js is to call
set_remote_log_functions(null) before anything publishes EV_ON_CLOSE.
That call clears the hooks and resets them to the console (see
helpers.js). The memory note “Remote-log unwire order” records the
incident.
9. The logcenter yuno¶
yunos/c/logcenter/ collects the logs that every yuno on the host, or on the
LAN, ships over UDP. It is not enabled by default. A yuno ships to UDP
only if its config lists a udp handler.
How it listens¶
UDP server (
c_gss_udp_s) onudp://127.0.0.1:1992by default (c_logcenter.c).Wire format:
<priority-digit><8hex-seq><json-payload><8hex-crc>, fragmented perudp_frame_size(default 1500,log_udp_handler.c).
What it does on receipt¶
In c_logcenter.c:
ac_on_message()parses each packet.Writes the JSON record to its own rotatory file
W.logviawrite2logs()/_log_bf(). Default size cap 600 MB (max_rotatoryfile_size, in megabytes).Updates in-memory counters per severity /
msgset/msg(do_log_stats(), c_logcenter.c:895).
What it exposes¶
Commands (c_logcenter.c):
| Command | Effect |
|---|---|
display-summary | Print the in-memory counters: alerts, criticals, errors, warnings. |
send-summary | Email the same summary (used as a daily/weekly batch). |
search | Search the stored log file for matching lines. |
tail | Last N lines of the centralized log. |
reset-counters | Zero the in-memory counters. |
Use it like any other yuno. Target it by yuno_role=logcenter, because the
numeric id of the yuno changes with the realm. command-yuno implies the
default service:
# rollup counters (Alert/Critical/Error/Warning/Info + Connect/Disconnect breakdown)
ycommand -c 'command-yuno yuno_role=logcenter command=display-summary'
# last N log lines (default ~100; can pass lines=N)
ycommand -c 'command-yuno yuno_role=logcenter command=tail lines=200'
# substring search (parameter is text=, not match=); maxcount caps the hits
ycommand -c 'command-yuno yuno_role=logcenter command=search text="EV_ON_CLOSE" maxcount=20'
# wipe the rollup counters — useful before reproducing an issue so the
# next display-summary only shows the new run
ycommand -c 'command-yuno yuno_role=logcenter command=reset-counters'Three more commands are useful (c_logcenter.c):
send-summary, enable-send-summary and disable-send-summary control the
email rollup. restart-yuneta-on-queue-alarm is the auto-recovery hook for a
UDP queue that floods.
Per-yuno vs centralized — when to use each¶
Per-yuno tail when you know which yuno is misbehaving and want raw control over
grep.<yuno dir>/logs/<file>.logis full fidelity.logcenter when you need correlation across multiple yunos, or when a yuno crashes too fast to read its own file, or for the rollup counters / email summaries.
Both can run at the same time. The file handler writes locally, and the UDP handler ships to logcenter in parallel. They are not exclusive.
10. Sharp edges¶
10.1 Traces accumulate — disable them when you finish¶
A forgotten set-global-trace level=machine set=1 survives a restart
(see §4 Persistence). The logs then grow without limit. Always pair the
enable and the disable in the same operational session. Before you leave, make
sure that the state is correct with get-global-trace.
10.2 Persistence asymmetry¶
| Scope | Persists across restart? |
|---|---|
global | Yes (via trace_levels attr) |
gclass | No |
gobj | No |
no_trace | No (all flavours) |
If you persisted a global level by mistake, clear it explicitly with
set-global-trace level=<name> set=0. To delete the file does not help,
because the value is in the treedb config of the yuno.
10.3 ievent_gate_stack is only on inter-event hops¶
A direct C function call between gobjs in the same yuno does not carry the stack, because there is no metadata to attach. The correlation exists only across yuno boundaries. Plan your traces for this limit.
10.4 LOG_AUDIT lines have no standard header¶
glogger.c writes the audit lines raw. A line filter that expects
the timestamp prefix misses them. When you look for operator actions, read the
audit file directly.
10.5 UDP can drop¶
UDP logs are not reliable. Under a burst, for example a machine trace that is fully on, the kernel buffer can overflow, and logcenter loses lines. It gives no message. The local file handler drops nothing, so trust the local file when you are not sure.
10.6 ev_kw is enormous¶
set-global-trace level=ev_kw set=1 writes the full kw JSON payload of every
event to the log. It is useful on a single-shot test. It is ruinous on a busy
service. If you need it narrowly, combine it with a machine trace that is
scoped to one gclass.
10.7 SPA dev-panel teardown order¶
set_remote_log_functions(null) MUST come before
do_disconnect / destroy_shell. See §8 and memory
feedback_remote_log_unwire_order.
10.8 Deep tracing has no ycommand switch¶
gobj_set_deep_tracing() is available only in C (gobj.c). If a yuno generates
traces that you cannot configure, look for a gobj_set_deep_tracing call that
someone left in its mt_create.
11. Operational recipes¶
11.1 Watch what a yuno does¶
YUNO=<id>
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-global-trace level=machine set=1"
tail -F /yuneta/realms/*/*/*^$YUNO/logs/*.log | grep -a '"msg":'
# reproduce
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-global-trace level=machine set=0"11.2 Watch traffic on a TCP/HTTP service¶
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-gclass-trace gclass=C_TCP_S level=traffic set=1"
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-gclass-trace gclass=C_PROT_HTTP_SR level=traffic set=1"
tail -F /yuneta/realms/*/*/*^$YUNO/logs/*.log
# … done …
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-gclass-trace gclass=C_TCP_S level=traffic set=0"
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-gclass-trace gclass=C_PROT_HTTP_SR level=traffic set=0"11.3 Watch all logs from this host in one place¶
Enable logcenter. Add this to the config JSON of every yuno:
"daemon_log_handlers": {
"to_udp": { "handler_type": "udp", "url": "udp://127.0.0.1:1992", "handler_options": 255 }
}Then:
ycommand -c 'command-yuno yuno_role=logcenter command=tail lines=500'
ycommand -c 'command-yuno yuno_role=logcenter command=search text="<keyword>" maxcount=50'
ycommand -c 'command-yuno yuno_role=logcenter command=display-summary'
ycommand -c 'command-yuno yuno_role=logcenter command=reset-counters' # wipe the rollup11.4 Follow one request end-to-end¶
See §7. The set of commands is set-gclass-trace ... traffic,
set-global-trace ... machine, set-global-trace ... fs and
set-gclass-trace C_IEVENT_SRV ... ievents2. Do not forget to disable
everything afterwards.
11.5 Capture an FSM bug in one gobj only¶
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-gobj-trace gobj=<short_name> level=machine set=1"
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-gobj-trace gobj=<short_name> level=ev_kw set=1"
tail -F /yuneta/realms/*/*/*^$YUNO/logs/*.log | grep -a '<short_name>'
# … done …
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-gobj-trace gobj=<short_name> level=machine set=0"
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-gobj-trace gobj=<short_name> level=ev_kw set=0"The disable step is not optional. These traces are persisted: a restart
does not clear them, it re-applies them. A forgotten ev_kw on a busy gobj
then writes for as long as the yuno lives.
11.6 Spot the canonical “Event NOT DEFINED in state” error¶
That single string is the most common FSM failure. The parent FSM did not
declare an event that a child publishes (see CLAUDE.md “CHILD vs SERVICE”).
The framework logs it at LOG_ERR, and the trace settings do not change that,
so:
grep -a '"msg":"Event NOT DEFINED in state"' /yuneta/realms/*/*/*/logs/*.logThis command works on any host, and it needs no trace.
11.7 Chase a leak to the line that allocated it¶
print_track_mem() runs at the end of every orderly shutdown of a build
with CONFIG_DEBUG_TRACK_MEMORY, and prints one mem-not-free line per block
still busy. It returns at once when nothing is busy, so no output means no
leak. Note where the report comes out: rotatory_end() runs before it, so
the yuno’s own FILE handler is already closed — the report reaches stdout and
the UDP handler (the logcenter), never logs/*.log.
Read it in three steps.
1. The size says WHAT the block is. track_mem_t costs 56 bytes and the
reported size includes it. Subtract it and the jansson structures appear:
| reported | real | what it is |
|---|---|---|
| 96 (+ 120) | 40 (+ 64) | a json_array_t and its initial 8-slot table |
| 88 (+ n) | 32 (+ n) | a json_string_t and its strlen+1 buffer |
2. The ref says WHEN it was ALLOCATED — not when it leaked. It is a
global allocation counter, so it dates the block. A block dated during the
treedb load can perfectly well be held by something built minutes later.
Equal deltas between blocks mean one turn of a loop each.
3. Ask for the bytes, and then for the stack. Two environment variables, read by any yuno:
# what the leaked blocks HOLD: 64 printable bytes each. The text blocks are
# the ones that name the leak -- a leaked string carries its characters.
YUNETA_TRACK_MEM_DUMP=1 <yuno> --config-file='[...]'
# photograph every allocation whose ref falls in a WINDOW, narrowed by size:
# each one logs its stack, and the stack is the allocation site.
YUNETA_TRACK_MEM=493000-498000:96,148,150 <yuno> --config-file='[...]'The window is what makes this usable. memory_check_list[] in the yuno’s
main.c needs the exact ref or the exact size: a ref cannot be prepared
in advance — it moves a few hundred between two runs of the same yuno — and a
size alone catches tens of thousands of blocks in a yuno that starts by
loading a database. Aim the window at what the previous run’s report showed,
and add the sizes of step 1. The header line of the report prints the window
it used, so a malformed value shows up as window_max: 0.
Worked example (db_history_ce, 2026-09-20): sixteen blocks read as six json
arrays plus two strings of 91 and 93 characters; YUNETA_TRACK_MEM_DUMP=1
printed the text of the strings, which named the field (observaciones of six
places nodes); the window gave the stacks, which put the allocation in the
treedb load — and that was the trap, because the holder was a configuration
built afterwards. What closed it was reproducing the leak in an isolated copy
of the yuno (see below) with a single call to the command that builds it.
Reproduce it off the live yuno. Copy the store under a scratch root, copy
the yuno’s config layers, point environment.work_dir at the scratch root,
add "yuno": {"autoplay": true} — nothing plays a yuno without the agent —
move the listening ports, and lower io_uring_entries (a second yuno of the
same size fails io_uring_queue_init_params() with ENOMEM). Validate the
copy before believing a negative: plant a gbmem_malloc(1234) in main()
and check the report prints it. A copy that does not leak is then a fact
about the difference between it and production, and that difference is the
answer.