Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Debugging a yuno

This document covers what you do when a yuno does not behave correctly. It explains how to enable the traces that show you what happens, where the output goes, and how to follow one message through several yunos. It also explains the part that the centralized log aggregator (logcenter) plays.

This document is the companion to YUNO_LIFECYCLE.md. That one covers how the agent manages yunos. This one covers how to look inside them.


1. Mental model

Three observation layers are independent. A confusion between them is the first source of frustration:

LayerQuestion it answersHow you turn it on
Log (severity)“Did something bad happen?”Always on. Filter by severity in the log file.
Trace (categories)“What was the system doing a moment ago?”set-global-trace / set-gclass-trace / set-gobj-trace — off by default.
Audit“What commands did operators run on this yuno?”Always written when use_audit_command_file=true.

You configure four destinations per yuno, with daemon_log_handlers in the yuno config JSON:

                        ┌──────────────────────────┐
       severity logs    │     file handler         │  → <yuno dir>/logs/<mask>.log
       + traces  ──────►│     (rotatory, ~8 MB)    │
                        └──────────────────────────┘
                        ┌──────────────────────────┐
                        │     udp handler          │  → udp://host:port
                        │     (default :1992)      │  → typically the logcenter yuno
                        └──────────────────────────┘
                        ┌──────────────────────────┐
                        │     stdout (console mode)│  → terminal when not daemonised
                        └──────────────────────────┘
                        ┌──────────────────────────┐
                        │     remote_log over      │
                        │     ievent / websocket   │  → SPA "dev panel" (live viewer)
                        └──────────────────────────┘

One log line can go to all four destinations at the same time. No destination is “the” log. They are different sinks.


2. Severity levels (gobj_log_*)

These are the calls every gclass uses to record events. They are not traces — they fire regardless of trace settings. The six public ones are defined in kernel/c/gobj-c/src/glogger.c:

glogger.c declares two more channels that are not syslog channels:

Per-yuno HARD RULE (see CLAUDE.md): every error-return path calls gobj_log_error or carries an // Error already logged comment. If you cannot find the error in the log, that yuno has a bug. The log did not lose it.


3. Trace categories

A trace is the running commentary that the framework can emit. It is off by default. It is noisy, so enable it only when you need it, and disable it when you finish.

3.1 Global trace levels

Defined in s_global_trace_level[16] at kernel/c/gobj-c/src/gobj.c:

BitNameEmits when
0machineEvery FSM event dispatch + every state change. The big one. See §6.
1create_deletegobj created / destroyed
2create_delete2Same as above, plus the kw payload
3subscriptionsgobj_subscribe_event / gobj_unsubscribe_event
4start_stopgobj_start / gobj_stop
5ev_kwDump the kw JSON payload on every event dispatch (huge volume)
6authzsAuthorization checks
7statesState changes (subset of machine)
8gbuffersgbuffer alloc / free / realloc
9timerOne-shot timer fires
10fsFilesystem ops — including timeranger2 appends
11liburingio_uring submit / complete
12timer_periodicPeriodic timer fires (separate from timer to avoid spam)
13liburing_timerio_uring-backed timers
14commandsgobj_command invocations

These are global bits. When you enable one, it affects every gobj in the yuno.

3.2 Per-gclass trace levels

Each gclass declares its own up-to-16 levels in s_user_trace_level[16]. Example: c_tcp_s.c

enum {
    TRACE_LISTEN        = 0x0001,
    TRACE_NOT_ACCEPTED  = 0x0002,
    TRACE_ACCEPTED      = 0x0004,
    TRACE_TLS           = 0x0008,
};
PRIVATE const trace_level_t s_user_trace_level[16] = {
    {"listen",          "Trace listen"},
    {"not-accepted",    "Trace not accepted connections"},
    {"accepted",        "Trace accepted connections"},
    {"tls",             "Trace tls"},
    {0, 0},
};

The names are gclass-specific. Common ones across runtime gclasses:

To see what a gclass offers, run get-gclass-trace gclass=<X> (see §4).

3.3 Per-gobj trace levels

These levels are the same as the per-gclass levels, but they are scoped to one gobj instance. They are useful when you have ten TCP connections and you want the trace of one connection. API: gobj_set_gobj_trace() at kernel/c/gobj-c/src/gobj.c:11256.

3.4 The no_trace parallel system

For every “set trace” command there is a “set no-trace” counterpart. The framework subtracts the no-trace mask from the effective trace mask. So you can enable a noisy level globally, then silence it on specific gclasses or gobjs. Functions: gobj_set_global_no_trace() at gobj.c:11711, gobj_set_gclass_no_trace() at gobj.c:11617, gobj_set_gobj_no_trace() at gobj.c:11746.

3.5 Deep trace mode

gobj_set_deep_tracing(level) enables all traces, and the masks do not apply. There is no ycommand for it. It is available only in the C API, and the framework uses it internally for emergency dumps. Do not use it unless you can accept the volume.


4. Turning traces on and off

All commands go to the yuno itself, addressed to its __yuno__ service. Handlers in kernel/c/root-linux/src/c_yuno.c:

# discover what a gclass offers
ycommand -c 'command-yuno id=<yuno> service=__yuno__ command=get-gclass-trace gclass=C_TCP_S'

# enable / disable
ycommand -c 'command-yuno id=<yuno> service=__yuno__ command=set-global-trace level=machine set=1'
ycommand -c 'command-yuno id=<yuno> service=__yuno__ command=set-gclass-trace gclass=C_TCP_S level=traffic set=1'
ycommand -c 'command-yuno id=<yuno> service=__yuno__ command=set-gobj-trace gobj=<short_name> level=machine set=1'

# silence (no_trace) — per gclass or per gobj
ycommand -c 'command-yuno id=<yuno> service=__yuno__ command=set-gclass-no-trace gclass=C_TIMER level=periodic set=1'
ycommand -c 'command-yuno id=<yuno> service=__yuno__ command=set-gobj-no-trace gobj=<short_name> level=machine set=1'

# inspect current state
ycommand -c 'command-yuno id=<yuno> service=__yuno__ command=get-global-trace'
ycommand -c 'command-yuno id=<yuno> service=__yuno__ command=get-gclass-trace gclass=C_TCP_S'
ycommand -c 'command-yuno id=<yuno> service=__yuno__ command=get-gobj-trace gobj=<short_name>'

The short form in CLAUDE.md, ycommand -c 'command-yuno id=<id> service=__yuno__ command=…', is exactly this. The shorter form ycommand -c 'set-global-trace …' sends command-yuno to the yuno that is registered as the default yuno.

set-global-no-trace silences a global level for every gobj. Each yuno’s main.c sets its defaults this way before it creates the yuno, almost always as gobj_set_global_no_trace("timer_periodic", TRUE). That is why a global machine trace does not drown in timer ticks. The command can undo such a default, and the change persists (see Persistence):

# see the periodic timer event, and keep seeing it after a restart
ycommand -c 'command-yuno id=<yuno> service=__yuno__ command=set-global-no-trace level=timer_periodic set=0'
ycommand -c 'command-yuno id=<yuno> service=__yuno__ command=set-global-trace level=timer_periodic set=1'

Persistence

Every trace set from the control plane is persisted. None of them are live-only. set-global-trace, set-gclass-trace and set-gobj-trace all end in gobj_save_persistent_attrs() on the yuno’s trace_levels attribute, and their no-trace counterparts on no_trace_levels:

CommandSaverKey in the attr
set-global-no-tracesave_global_no_trace__global_no_trace__ (in no_trace_levels)
set-global-tracesave_global_trace__global_trace__
set-gclass-tracesave_user_tracethe gclass name
set-gobj-tracesave_user_tracethe gobj name
set-gclass-no-tracesave_user_no_tracethe gclass name
set-gobj-no-tracesave_user_no_tracethe gobj name

They are re-applied on the next start: set_user_gclass_traces() and set_user_gclass_no_traces() run from mt_create, set_user_gobj_traces() right after the children are built.

A saved scope REPLACES what is in force. The global scope, the global no-trace scope and the scope of each gclass are saved whole, from the levels in force after the command, and an empty scope is saved as []. At the next start up, a scope that is in the attr replaces what main() set before it created the yuno. A scope that was never saved keeps the default of main(). So a default that the user turned off stays off:

"no_trace_levels": {
    "__global_no_trace__": [],
    "C_TIMER": ["machine"]
}

With this attr, timer_periodic is not silenced after a restart, although main.c silences it, and C_TIMER keeps its machine no-trace. (A scope saved by an older release is the list of levels that were set one by one; it now replaces the defaults of main() too.) A gobj-name key, which reset-all-traces gobj=… writes, keeps its old level-by-level list.

CAUTION: a forgotten set-gclass-trace gclass=C_TCP_S level=traffic set=1 survives a restart exactly like a global one, and it fills your disk. It gives no message first. Always pair the enable and the disable in the same session.

An entry whose key names a gclass that no longer exists is skipped without a log at start up (c_yuno.c:4992, a deliberate exception to the no-silent-errors rule: a mistyped gclass name is the common case). So a trace that does not turn on after a restart usually means a typo in the persisted attr. Read it with list-persistent-attrs, and clear it with remove-persistent-attrs.

4.1 When the yuno never reaches the agent (--global-trace)

Every command above travels over the yuno’s control channel to the agent. So none of them work for the failure that most needs a trace: a yuno that dies, hangs or fails before that channel is ready. ycommand cannot reach it, and list-yunos reports running=false even while the process is alive.

Since 7.8.2 the levels can be armed on the command line instead:

# one level, or several — repeatable and comma-separated
auth_bff --config-file='[...]' --global-trace=machine
auth_bff --config-file='[...]' --global-trace=machine,create_delete,start_stop

# what levels exist
auth_bff --global-trace=list

The framework applies them after it registers every gclass, and before the first service starts, so they cover start up itself. It applies them again right after it creates the yuno, because the yuno replaces the global scope with the persisted one: what you ask on the command line wins. A later trace command saves the global levels in force, and those include the ones from the command line. An unknown level stops the yuno with a message that points at list. The yuno does not ignore it.

To reproduce a yuno that the agent launches, take its command line from running-bin id=<id> or running-keys id=<id>. You can also use the script that the agent writes at /yuneta/realms/<realm>/<yuno>/bin/<role>^<id>.sh. Then append the flag.

Two older methods, and their limits:

4.2 Two warnings that arrive without being asked for

Some failures below the framework cannot wait for a trace, so they report themselves:

A dead first nameserver costs ~6 s (A + AAAA timeouts) on every lookup, so a yuno that opens many channels can spend minutes in start up. The resolver caches answers since 7.8.2, which limits the cost to the first lookup. But the correction belongs in the node’s /etc/resolv.conf.


5. Reading the logs

5.1 File paths

Per-yuno log file, built by yuneta_log_file() at kernel/c/root-linux/src/yunetas_environment.c: <work_dir>/<domain_dir>/logs/<filename_mask>. work_dir is /yuneta. The agent gives each yuno that it runs its own domain_dir (build_yuno_private_domain() in c_agent.c):

/yuneta/realms/<realm_owner>/<realm_name>.<realm_role>.<realm_env>/<role>^<id>/logs/<filename_mask>

For example, /yuneta/realms/artgins/artgins.yunetacontrol.com/controlcenter^1996/logs/controlcenter-4.log. The agent itself and the utilities run by hand have a fixed domain_dir in their main.c: /yuneta/realms/agent/agent/logs/ for yuneta_agent, /yuneta/realms/agent/<utility>/logs/ for ycommand, ybatch and the others. There is no /yuneta/logs/ directory.

The mask is the value that you set in daemon_log_handlers.<handler>.filename_mask (see §5.4). By convention it is <role>-W.log, where the day of the week (1 = Sunday … 7) replaces the W: on a Wednesday the yuno writes <role>-4.log.

Active log discovery:

ls -lt /yuneta/realms/*/*/*^<id>/logs/
tail -f /yuneta/realms/*/*/*^<id>/logs/<latest>.log | grep -a "keyword"

5.2 Log line format

Every log record is a JSON object built in glogger.c. Fields added automatically by discover() at glogger.c:1326:

FieldSource
timestampcurrent_timestamp()
priorityLOG_ERR / LOG_WARNING / …
node_uuidhost node identity
processyuno binary name
hostnamefrom gethostname
pidprocess id
gclassthe gclass that emitted the line
gobj_namethe gobj instance name
statecurrent FSM state of that gobj
gobj_full_namedotted path (only if gobj_full_name trace is on)
idsequence id
msgset, msgthe "msgset","msg" pair every gobj_log_* call passes
any key,valueextra fields the caller passed

Searching is JSON-friendly:

grep -a '"priority":3' <yuno dir>/logs/<file>.log       # all errors
grep -a '"gclass":"C_TCP_S"' <yuno dir>/logs/<file>.log  # one gclass
grep -a '"msg":"Event NOT DEFINED in state"' …               # the canonical FSM bug

5.3 Rotation

The rotatory library makes the file name from the mask and the local date, so the file changes at midnight (the first record after it). With the W mask, the first record of a new day, and the first record after the yuno starts, empties the file of the same week day when it was last written before today: that is last week’s file. So a yuno keeps 7 days of log, and the mask is the retention. A file written today is appended to, so a restart keeps it, and a clock set back across midnight empties nothing. A mask with no date letter (a fixed name such as logcenter.log) is never emptied by a date. The table of which masks empty a file is in File names and rotation.

The library also rotates the file when it becomes larger than a size threshold (default 8 MB, counted in bytes, configurable via max_megas_rotatoryfile_size, entry_point.c): it renames the file to <name>.OLD (a previous .OLD is removed) and starts the file again. There is no cron. Both rotations happen on the next write. The file is checked once for each record, never between the pieces of a record, so a record is never split between two files. A log file removed by hand is created again by the next record. At a new day the size of the file of the day before is not checked: that file is left as it is, even over the limit, and its .OLD stays. (In 7.25.4, when the last piece of a day took its file over the limit, the first record of the next day renamed that file to .OLD and removed the .OLD of that day.)

A log file RENAMED by another program (a logrotate with its default create mode) is not noticed: the yuno goes on writing into the renamed file until its next new file (up to 7.25.4 it was noticed). To rotate a yuno log from outside, copy and truncate it (logrotate with copytruncate), or remove it:

cp /yuneta/realms/agent/agent/logs/yuneta_agent-4.log /tmp/ \
    && truncate -s 0 /yuneta/realms/agent/agent/logs/yuneta_agent-4.log

A full disk stops only that file. Every 100 records the handle checks the free space of its own disk; below min_free_disk_percentage (default 10%) it drops its records, and it writes again by itself when the space is back. It prints one line to stdout and syslog when it stops and one when it resumes (up to 7.25.4 one full disk stopped every file log of the process until it was restarted):

rotatory(): stop logging to '/yuneta/realms/agent/agent/logs/yuneta_agent-4.log' because full disk: 9% free (<10%)
rotatory(): logging to '/yuneta/realms/agent/agent/logs/yuneta_agent-4.log' again: 12% free (>=10%), 5210 records were dropped

See a full disk in the rotatory page.

With the defaults, one yuno uses at most 7 × 2 × 8 MB = 112 MB of log (each piece ends with the record that took it over 8 MB). Up to 7.25.4 the size was counted in whole megabytes, so a piece rotated only at 9 MB: 126 MB.

5.4 Where to configure handlers

In the yuno’s config JSON, under environment.daemon_log_handlers (or console_log_handlers in non-daemon mode), parsed at kernel/c/root-linux/src/entry_point.c:

"environment": {
    "daemon_log_handlers": {
        "to_file": {
            "handler_type": "file",
            "filename_mask": "mqtt_broker-W.log",
            "handler_options": 255
        },
        "to_udp": {
            "handler_type": "udp",
            "url": "udp://127.0.0.1:1992",
            "handler_options": 255
        }
    }
}

handler_options is a bitmask of LOG_HND_OPT_* (glogger.h) that selects which severities the handler accepts. 255 accepts all of them. If you clear bits, the handler drops DEBUG, INFO, AUDIT and the other severities.

To add or remove handlers at run time, use the add-log-handler and del-log-handler commands of c_yuno.c.

5.5 The agent’s audit files

yuneta_agent writes every command that it runs to an audit file, one JSON record for each command (use_audit_command_file, on by default). The console writes (write-tty, one command for each keystroke) are the exception: they make one record for each burst (see Console writes). The files are in the audit/ directory of the agent realm:

/yuneta/realms/agent/agent/audit/266-23_09_2026.log        # ZZZ-DD_MM_CCYY.log
/yuneta/realms/agent/agent/audit/266-23_09_2026.log.OLD.1  # the first 500 MB of a big day
/yuneta/realms/agent/agent/audit/266-23_09_2026.log.OLD.2  # the next 500 MB

What a record holds

The record is built by audit_record_build() (yunos/c/yuno_agent/src/audit_record.c). There are two forms, and a third one for the console writes.

A read-only command gets only the command, the date and the user:

{"command":"list-yunos","date":"2026-09-24T10:00:00.000000000+0200","user":"claudia@artgins.com"}

The command tables do not say which commands are read-only, so the list is by name: the prefixes list-, view-, get-, info-, dir-, and help, authzs, ping, node-uuid, top, top-services, services, stats, stats-agent, stats-yuno, authzs-yuno, treedbs, treedb-info, topics, desc, descs, system-schema, schema-file, saved-schema, diff-schema, jtree, nodes, node, instances, hooks, links, parents, children, pkey2s, snaps, snap-content, print-role, print-tranger, check-json, check-realm, cert-expiry-status, cert-sync-status, global-variables, running-keys, running-bin, users, accesses, roles, user-roles, user-authzs.

A command that has a __reset__ value (in the command text or in the kw) is not read-only: stats=__reset__ sets the counters of a yuno to zero. So stats-yuno id=gate_mqtts stats=__reset__ (the reset button of gui_agent) gets the full record, with its source, also when the __reset__ comes in the kw.

command-yuno and command-agent are judged by the command that they carry, and the record names it: "command":"command-yuno command=view-attrs". The carried command is taken from the same place as the handler takes it: the last command= of the command text, else command of the kw. So a kw command=list-yunos with a text command='delete-node …' runs delete-node and is recorded as delete-node, with the full record.

The key must be exactly command. The command parser matches a key of the text in any case, but it stores the value under the key as it was typed, and the handler reads command only. So COMMAND=list-yunos is not the command that runs: with a kw {"command": "delete-yuno id=gate"}, the command

command-yuno id=gate COMMAND=list-yunos

runs delete-yuno, and it is recorded with the full record (the kw, the source). A key command in another case, in the text or in the kw, always gives the full record: it is never taken as read-only. The same rule gives the console of a write-tty: the handler reads name, so NAME=decoy does not change the console of the record.

These commands keep the full record: read-file, read-json and read-binary-file (they read files of the node), check-user-pwd, and anything that opens something (open-list, open-treedb, …).

Every other command gets the command, the date, the user, the source and the parameters:

{"command":"install-binary id=auth_bff content64='<33554432 bytes sha256:401b36b9e4f91e967e815f96fd293cb384222d3e73e1b9fcf759d5ecfadfdbf8>'",
 "date":"2026-09-24T10:00:00.000000000+0200",
 "user":"yuneta",
 "source":{"hops":[{"role":"ycommand","yuno":"","service":"ycommand",
                    "user":"yuneta","host":"dev-laptop"}]},
 "kw":{"__username__":"yuneta"}}

The same command sent from gui_agent through the controlcenter has two hops, nearest first:

"source":{"console_purpose":"statnodes",
          "hops":[{"role":"controlcenter","yuno":"artgins.com","service":"top-16",
                   "user":"yuneta_agent@artgins.com","host":"artgins"},
                  {"role":"gui_agent","yuno":"gui_agent_yuno","service":"agent_link",
                   "user":"claudia@artgins.com","host":"544f1345-65a6-455c-aa94-62b6b020b5c5"}]}

Up to 7.25.4 the record was the command, the date and the WHOLE kw. The sizes, measured with the same commands:

CommandUp to 7.25.4Now
install-binary of a 32 MB yuno (content64 in the command text, as ycommand sends it)134,218,673 bytes (the base64 three times)541 bytes
run-yuno through the controlcenter891 bytes489 bytes
list-yunos through the controlcenter840 bytes97 bytes

On wattyzer, a deploy day wrote 0.6–1.2 GB of audit (5 to 9 binaries) and a normal day 2 KB to 2.6 MB. With this format a deploy day writes a few KB.

Console writes

ycommand, ycli and gui_agent send one write-tty command for each keystroke typed into an agent console, with the keystroke in content64. The audit keeps only the fact: who, when, which console, how many writes and bytes. Nothing of what was typed is written, not even a hash: the sha256 of one byte can be read back with a table of 256 entries, so a hash would give the typed text (and a typed password) back.

The writes of one user (the same user and the same source) into one console make a burst. A burst lasts 60 seconds from its first write.

{"command":"write-tty","date":"2026-09-24T10:00:00.1+0200","user":"claudia@artgins.com",
 "console":"console-1","writes":1,"bytes":1,"source":{…}}
{"command":"write-tty","date":"2026-09-24T10:00:00.4+0200","user":"claudia@artgins.com",
 "console":"console-1","writes":212,"bytes":230,"until":"2026-09-24T10:00:58.9+0200","source":{…}}

The records of a burst do not overlap, so the sum of their writes and bytes is the whole burst. A burst ends:

Up to 7.25.4 each keystroke wrote the whole kw, with the keystroke in base64 (about 920 bytes each). Now 1000 keystrokes in one burst write two records, about 950 bytes. A write-tty carried by command-agent gets its full record, with content64 as <N bytes> only.

Rotation and retention

The mask ZZZ-DD_MM_CCYY.log makes a new name every day. When the file of the day crosses max_megas_audit_file, it is renamed to the first free .OLD.<n> (.OLD.1, .OLD.2, …) and a new file begins. No piece of a day is removed at a rotation: the audit rotatory uses rotatory_keep_all_old_files(). Up to 7.25.4 each rotation removed the previous .OLD, so a day that crossed the limit twice lost its first part (on wattyzer, the mornings of 22 and 23 September 2026).

If the directory refuses the rename (chattr +a on it, a read-only bind, a MAC denial), nothing is removed: the file of the day is kept and grows over the limit. The agent prints one line (stdout and syslog) and tries the rename again after 60 seconds of real time (the monotonic clock: a clock set back or forward does not move the retry), or at the next day, not at every command. It prints one more line when a rename works again:

_rotatory(): Cannot rename '/yuneta/realms/agent/agent/audit/267-24_09_2026.log' to '/yuneta/realms/agent/agent/audit/267-24_09_2026.log.OLD.1', Permission denied, the file is kept and grows, the rename is tried again every minute
_rotatory(): the size rotation of '/yuneta/realms/agent/agent/audit/267-24_09_2026.log' works again

The attribute audit_keep_days (default 7) is the retention. The agent removes the audit files older than that number of days:

The agent opens its audit with exit_on_fail (the last parameter of rotatory_open()), and each yuno opens its file log the same way: the process does not start if that file cannot be opened. This applies to that first open only. A failure after the start never stops the agent or the yuno. This applies to a new day, a file removed from the directory, the open after a failed write, and a truncate. The agent prints one line (stdout and syslog), and the command runs without its audit record. The next command tries the open again, with no line while it fails. It prints one line when the open works again, and the retention runs then (see exit_on_fail is for the open):

_rotatory(): Cannot create '/yuneta/realms/agent/agent/audit/269-26_09_2026.log' file, Too many open files
_rotatory(): '/yuneta/realms/agent/agent/audit/269-26_09_2026.log' is open again

Up to 7.25.4 the first failure after the start exited the agent, and the retention of that day did not run.

The second sweep runs inside the write of the first record of the new file, before that record is written: so it is on the write path of that one command, once a day (or once for each size rotation). The same file opened again (after a write that failed, or when the file was removed) is not a new file, and it runs no sweep. It reads the directory once. It removes only regular files with the name shape of the mask (and their .OLD / .OLD.<n>), never a file of the current day, never a symbolic link, never another file in the directory (rotatory_remove_old_files()). Each sweep that removes something writes one INFO line to the agent log:

{"msg": "Old audit files removed", "audit_keep_days": 7, "removed": 3, "megas": 961,
 "current_file": "/yuneta/realms/agent/agent/audit/266-23_09_2026.log",
 "files": ["258-15_09_2026.log", "258-15_09_2026.log.OLD.1", "258-15_09_2026.log.OLD.2"]}
AttributeDefaultMeaning
use_audit_command_file1Write the audit files.
max_megas_audit_file500Size of one piece of a day, in MB. A bigger day continues in .OLD.<n> pieces.
audit_keep_days7Days of audit files kept. 0 keeps all (the behaviour up to 7.25.4).
min_free_disk_percentage20Stop writing the audit when the disk has less free space (checked every 100 records), and write it again when the space is back. A new day still runs the retention while the disk is full: the retention is what frees the space. If the new file of that day cannot be opened, the open is tried again every 100 records until it works, and then the retention runs.

With the new record format the directory is a few MB a week. The retention still bounds it by days, whatever a day writes.

To keep 30 days, set the attribute in the agent config (/yuneta/agent/yuneta_agent.json) and restart the agent:

{
    "global": {
        "agent.audit_keep_days": 30
    }
}

To change it on a running agent (it applies at the next new audit file, and is lost at restart):

ycommand -S __yuno__ -c 'write-attr gobj=agent attribute=audit_keep_days value=30'
ycommand -S __yuno__ -c 'view-attrs gobj=agent attribute=audit_keep_days'

An existing large audit/ directory (up to 7.25.4 nothing was removed, and 19 GB and 90 GB were seen) needs no manual action. The first start of an agent with this version removes every audit file older than 7 days, logs one INFO line with the list, and keeps the last 7 days and today. On a slow disk this first sweep can take some seconds, once. To keep more, set audit_keep_days before that start.

A clock set back across midnight empties no audit file: the file of the day before is opened again and the records are appended to it. Up to 7.25.4 it was opened with "w" and emptied, and the file of today was emptied too when the clock went forward again (see the rotatory).

A tool that reads the audit must accept both formats: the files written before the upgrade have the whole kw (with __md_iev__ and the base64) and no user or source. It must also accept the write-tty records, which have console, writes, bytes and until and no kw, and <redacted> in place of a secret.


6. The FSM trace (machine)

This is the most useful trace for the behavior of a gobj. It is defined in glogger.c (trace_machine). Called from the event dispatcher in gobj.c:

Two output formats, switched by the integer variable trace_machine_format.

Format 1 — one line per transition. THE DEFAULT, on a node and in the browser alike (gobj.c: trace_machine_format = 1; // 0 legacy, 1 simpler; gobj-js followed at 7.13.7):

🔜 EV_RX_DATA !!c_tcp :open
🔄 EV_RX_DATA !!c_tcp :open from !!service_main
🔝🔝 EV_ON_MESSAGE c_prot_tcp4h^output-0 :wait_payload
🔝🔄 EV_ON_MESSAGE (EV_ON_MESSAGE) c_channel^output-0

Format 1 writes no return line and no state line: the transition line already carries the state it ran in.

Format 0 — legacy, three lines for one transition:

🔜 mach(!!c_tcp), st: :open, ev: EV_RX_DATA, from(!!service_main)
🔄 mach(!!c_tcp), st: :open, ev: EV_RX_DATA, from(c_tcp_s^server)
🔀🔀 mach(!!c_tcp), new st(:closed), prev st(:open)
<- mach(!!c_tcp), st: :closed, ev: EV_RX_DATA, ret: 0

Every line of either format is indented by its nesting depth, two spaces per level, so an event sent from inside another one’s action sits under it. That indentation is the only thing that says a transition happened during another — keep it when you render the line anywhere else.

!! before a name means that the gobj is not running at that moment. Two of them in a row are usually the bug.

Scoping the machine trace

The machine trace of a whole yuno gives too much output on anything larger than a toy test. You can make it narrow in two ways:

# only one gclass
ycommand -c 'command-yuno id=<yuno> service=__yuno__ command=set-gclass-trace gclass=C_TCP_S level=machine set=1'

# only one instance
ycommand -c 'command-yuno id=<yuno> service=__yuno__ command=set-gobj-trace gobj=<short_name> level=machine set=1'

Both are persisted and are re-applied on the next start, like every other trace set from the control plane (see Persistence). Disable them in the same session.

The same trace in the browser

This chapter is about a yuno on a node. But a browser SPA runs the same kernel, ported to JavaScript. Since @yuneta/gobj-js 7.9.5 the JS runtime has this level model, with the same names and the same bits. A habit that you learn here therefore transfers, and you can read two traces side by side.

gobj_set_global_trace("machine", true);          // the big one, same as above
gobj_set_gclass_trace("C_MY_VIEW", "machine", true);
gobj_set_gobj_no_trace(noisy_src, "machine", true);   // veto, by the SOURCE

set_log_callback((level, msg) => { ... });       // the trace arrives as `debug`

There is no ycommand on that side. The switch is the call above, and the output goes to the browser console. It goes to any other destination that set_log_callback() selects, and that is how the dev panel of gobj-ui shows the machine inside the app. doc.yuneta.io/navigation runs three demos with the panel connected to that callback. Read them if you want to see the lines before you write your own code.


7. Following a message end-to-end

Canonical request flow on a typical Yuneta service:

   external client
         │
         ▼
  ┌─────────────┐   gclass trace 'traffic'
  │   C_TCP_S   │   gobj_trace_dump_gbuf(gobj, gbuf, …)
  └──────┬──────┘
         │
         ▼
  ┌─────────────────┐   gclass trace 'traffic'
  │ C_PROT_HTTP_SR  │
  └────────┬────────┘
           │
           ▼
  ┌─────────────────┐   gclass trace 'ievents' / 'ievents2'
  │   C_IEVENT_SRV  │   trace_inter_event2(gobj, prefix, event, kw)
  └────────┬────────┘
           │
           ▼
  ┌─────────────────┐   global trace 'machine' lights up the FSM dispatch
  │  service gclass │   gclass-specific traces fire its custom emit points
  └────────┬────────┘
           │
           ▼
  ┌─────────────────┐   global trace 'fs'
  │   timeranger2   │   record append + rowid emitted
  │   (treedb)      │
  └────────┬────────┘
           │
           ▼  outbound publish
  ┌─────────────────┐   gclass trace 'ievents' / 'ievents2'
  │   C_IEVENT_SRV  │
  └────────┬────────┘
           │
           ▼
  ┌─────────────────┐   gclass-specific trace
  │  C_WEBSOCKET    │   gobj_trace_dump frames
  └────────┬────────┘
           │
           ▼
        SPA browser

The correlation id

Inter-event messages between yunos carry a metadata block named __md_iev__ inside the kw. Inside it is the ievent_gate_stack — a LIFO of hops, each entry: {src_yuno, src_service, dst_yuno, dst_service, user, host, …}.

To grep the same transaction across multiple yunos’ logs:

grep -a 'ievent_gate_stack' /yuneta/realms/*/*/*/logs/*.log | grep '<the user or src_yuno you care about>'

The framework propagates no automatic UUID for calls that are not ievents. A direct C function call has nothing to grep. The correlation is available only when the message crosses an ievent boundary.

Practical sequence to follow one HTTP request

YUNO=my_service_01

# 1. ingress + protocol
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-gclass-trace gclass=C_TCP_S        level=traffic set=1"
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-gclass-trace gclass=C_PROT_HTTP_SR level=traffic set=1"

# 2. internal FSM
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-global-trace level=machine set=1"

# 3. broker/topic write
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-global-trace level=fs set=1"

# 4. egress to SPA
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-gclass-trace gclass=C_IEVENT_SRV level=ievents  set=1"
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-gclass-trace gclass=C_IEVENT_SRV level=ievents2 set=1"

# trigger the request, capture the noise
tail -F /yuneta/realms/*/*/*^$YUNO/logs/*.log > /tmp/$YUNO.trace &
# … reproduce …
kill %1

# disable everything
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-gclass-trace gclass=C_TCP_S        level=traffic  set=0"
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-gclass-trace gclass=C_PROT_HTTP_SR level=traffic  set=0"
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-global-trace level=machine set=0"
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-global-trace level=fs      set=0"
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-gclass-trace gclass=C_IEVENT_SRV level=ievents  set=0"
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-gclass-trace gclass=C_IEVENT_SRV level=ievents2 set=0"

8. The SPA-side “dev panel” viewer

A SPA built on the JS gobj framework can connect to a yuno. Then it can display live what crosses the websocket: the same lines that you see in the local log file, plus the bodies of the ievent messages.

Wire-up

The trace buttons are the yuno’s commands

Each button of the window turns trace bits on or off with the trace commands of the JS C_YUNO, which has the commands and the attrs of the C one (set-global-trace, set-global-no-trace, set-gclass-trace, set-gclass-no-trace; trace_levels, no_trace_levels). The yuno persists them through the functions that the app gives to gobj_start_up(), and restores them in its mt_create with the same rule as C: a saved scope replaces the defaults of main.js.

ButtonWhat it sets
Automataglobal machine; a second click adds ev_kw
Creationglobal create_delete
Start / Stopglobal start_stop
Subscriptionsglobal subscriptions
TrafficC_IEVENT_CLI level ievents
Periodicglobal timer_periodic, and it clears the global no-trace of the same level
I18nnot a trace level: i18next’s debug switch, kept in the browser

The window never reads a message to decide whether to show it. The state of a button is read from the runtime, so it cannot disagree with what is traced.

Filtering on the SPA side

The window filters only what it shows: the direction (in / out / error) and a free text. To change what is traced, use the buttons, which change what the gobjs emit.

Teardown order — the recursion gotcha

When the websocket closes, ac_on_close (c_ievent_cli.js:897) fires EV_ON_CLOSE. set_remote_log_functions redirects the JS log_error and log_warning calls to the DOM callback. If it is still installed, the callback captures the warning that the teardown path emits. The callback changes the DOM, the change can fire more events, and those events log again. The result is an infinite recursion.

The correction at c_ievent_cli.js is to call set_remote_log_functions(null) before anything publishes EV_ON_CLOSE. That call clears the hooks and resets them to the console (see helpers.js). The memory note “Remote-log unwire order” records the incident.


9. The logcenter yuno

yunos/c/logcenter/ collects the logs that every yuno on the host, or on the LAN, ships over UDP. It is not enabled by default. A yuno ships to UDP only if its config lists a udp handler.

How it listens

What it does on receipt

In c_logcenter.c:

What it exposes

Commands (c_logcenter.c):

CommandEffect
display-summaryPrint the in-memory counters: alerts, criticals, errors, warnings.
send-summaryEmail the same summary (used as a daily/weekly batch).
searchSearch the stored log file for matching lines.
tailLast N lines of the centralized log.
reset-countersZero the in-memory counters.

Use it like any other yuno. Target it by yuno_role=logcenter, because the numeric id of the yuno changes with the realm. command-yuno implies the default service:

# rollup counters (Alert/Critical/Error/Warning/Info + Connect/Disconnect breakdown)
ycommand -c 'command-yuno yuno_role=logcenter command=display-summary'

# last N log lines (default ~100; can pass lines=N)
ycommand -c 'command-yuno yuno_role=logcenter command=tail lines=200'

# substring search (parameter is text=, not match=); maxcount caps the hits
ycommand -c 'command-yuno yuno_role=logcenter command=search text="EV_ON_CLOSE" maxcount=20'

# wipe the rollup counters — useful before reproducing an issue so the
# next display-summary only shows the new run
ycommand -c 'command-yuno yuno_role=logcenter command=reset-counters'

Three more commands are useful (c_logcenter.c): send-summary, enable-send-summary and disable-send-summary control the email rollup. restart-yuneta-on-queue-alarm is the auto-recovery hook for a UDP queue that floods.

Per-yuno vs centralized — when to use each

Both can run at the same time. The file handler writes locally, and the UDP handler ships to logcenter in parallel. They are not exclusive.


10. Sharp edges

10.1 Traces accumulate — disable them when you finish

A forgotten set-global-trace level=machine set=1 survives a restart (see §4 Persistence). The logs then grow without limit. Always pair the enable and the disable in the same operational session. Before you leave, make sure that the state is correct with get-global-trace.

10.2 Persistence asymmetry

ScopePersists across restart?
globalYes (via trace_levels attr)
gclassNo
gobjNo
no_traceNo (all flavours)

If you persisted a global level by mistake, clear it explicitly with set-global-trace level=<name> set=0. To delete the file does not help, because the value is in the treedb config of the yuno.

10.3 ievent_gate_stack is only on inter-event hops

A direct C function call between gobjs in the same yuno does not carry the stack, because there is no metadata to attach. The correlation exists only across yuno boundaries. Plan your traces for this limit.

10.4 LOG_AUDIT lines have no standard header

glogger.c writes the audit lines raw. A line filter that expects the timestamp prefix misses them. When you look for operator actions, read the audit file directly.

10.5 UDP can drop

UDP logs are not reliable. Under a burst, for example a machine trace that is fully on, the kernel buffer can overflow, and logcenter loses lines. It gives no message. The local file handler drops nothing, so trust the local file when you are not sure.

10.6 ev_kw is enormous

set-global-trace level=ev_kw set=1 writes the full kw JSON payload of every event to the log. It is useful on a single-shot test. It is ruinous on a busy service. If you need it narrowly, combine it with a machine trace that is scoped to one gclass.

10.7 SPA dev-panel teardown order

set_remote_log_functions(null) MUST come before do_disconnect / destroy_shell. See §8 and memory feedback_remote_log_unwire_order.

10.8 Deep tracing has no ycommand switch

gobj_set_deep_tracing() is available only in C (gobj.c). If a yuno generates traces that you cannot configure, look for a gobj_set_deep_tracing call that someone left in its mt_create.


11. Operational recipes

11.1 Watch what a yuno does

YUNO=<id>
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-global-trace level=machine set=1"
tail -F /yuneta/realms/*/*/*^$YUNO/logs/*.log | grep -a '"msg":'
# reproduce
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-global-trace level=machine set=0"

11.2 Watch traffic on a TCP/HTTP service

ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-gclass-trace gclass=C_TCP_S        level=traffic set=1"
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-gclass-trace gclass=C_PROT_HTTP_SR level=traffic set=1"
tail -F /yuneta/realms/*/*/*^$YUNO/logs/*.log
# … done …
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-gclass-trace gclass=C_TCP_S        level=traffic set=0"
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-gclass-trace gclass=C_PROT_HTTP_SR level=traffic set=0"

11.3 Watch all logs from this host in one place

Enable logcenter. Add this to the config JSON of every yuno:

"daemon_log_handlers": {
    "to_udp": { "handler_type": "udp", "url": "udp://127.0.0.1:1992", "handler_options": 255 }
}

Then:

ycommand -c 'command-yuno yuno_role=logcenter command=tail lines=500'
ycommand -c 'command-yuno yuno_role=logcenter command=search text="<keyword>" maxcount=50'
ycommand -c 'command-yuno yuno_role=logcenter command=display-summary'
ycommand -c 'command-yuno yuno_role=logcenter command=reset-counters'   # wipe the rollup

11.4 Follow one request end-to-end

See §7. The set of commands is set-gclass-trace ... traffic, set-global-trace ... machine, set-global-trace ... fs and set-gclass-trace C_IEVENT_SRV ... ievents2. Do not forget to disable everything afterwards.

11.5 Capture an FSM bug in one gobj only

ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-gobj-trace gobj=<short_name> level=machine set=1"
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-gobj-trace gobj=<short_name> level=ev_kw   set=1"
tail -F /yuneta/realms/*/*/*^$YUNO/logs/*.log | grep -a '<short_name>'
# … done …
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-gobj-trace gobj=<short_name> level=machine set=0"
ycommand -c "command-yuno id=$YUNO service=__yuno__ command=set-gobj-trace gobj=<short_name> level=ev_kw   set=0"

The disable step is not optional. These traces are persisted: a restart does not clear them, it re-applies them. A forgotten ev_kw on a busy gobj then writes for as long as the yuno lives.

11.6 Spot the canonical “Event NOT DEFINED in state” error

That single string is the most common FSM failure. The parent FSM did not declare an event that a child publishes (see CLAUDE.md “CHILD vs SERVICE”). The framework logs it at LOG_ERR, and the trace settings do not change that, so:

grep -a '"msg":"Event NOT DEFINED in state"' /yuneta/realms/*/*/*/logs/*.log

This command works on any host, and it needs no trace.

11.7 Chase a leak to the line that allocated it

print_track_mem() runs at the end of every orderly shutdown of a build with CONFIG_DEBUG_TRACK_MEMORY, and prints one mem-not-free line per block still busy. It returns at once when nothing is busy, so no output means no leak. Note where the report comes out: rotatory_end() runs before it, so the yuno’s own FILE handler is already closed — the report reaches stdout and the UDP handler (the logcenter), never logs/*.log.

Read it in three steps.

1. The size says WHAT the block is. track_mem_t costs 56 bytes and the reported size includes it. Subtract it and the jansson structures appear:

reportedrealwhat it is
96 (+ 120)40 (+ 64)a json_array_t and its initial 8-slot table
88 (+ n)32 (+ n)a json_string_t and its strlen+1 buffer

2. The ref says WHEN it was ALLOCATED — not when it leaked. It is a global allocation counter, so it dates the block. A block dated during the treedb load can perfectly well be held by something built minutes later. Equal deltas between blocks mean one turn of a loop each.

3. Ask for the bytes, and then for the stack. Two environment variables, read by any yuno:

# what the leaked blocks HOLD: 64 printable bytes each. The text blocks are
# the ones that name the leak -- a leaked string carries its characters.
YUNETA_TRACK_MEM_DUMP=1 <yuno> --config-file='[...]'

# photograph every allocation whose ref falls in a WINDOW, narrowed by size:
# each one logs its stack, and the stack is the allocation site.
YUNETA_TRACK_MEM=493000-498000:96,148,150 <yuno> --config-file='[...]'

The window is what makes this usable. memory_check_list[] in the yuno’s main.c needs the exact ref or the exact size: a ref cannot be prepared in advance — it moves a few hundred between two runs of the same yuno — and a size alone catches tens of thousands of blocks in a yuno that starts by loading a database. Aim the window at what the previous run’s report showed, and add the sizes of step 1. The header line of the report prints the window it used, so a malformed value shows up as window_max: 0.

Worked example (db_history_ce, 2026-09-20): sixteen blocks read as six json arrays plus two strings of 91 and 93 characters; YUNETA_TRACK_MEM_DUMP=1 printed the text of the strings, which named the field (observaciones of six places nodes); the window gave the stacks, which put the allocation in the treedb load — and that was the trap, because the holder was a configuration built afterwards. What closed it was reproducing the leak in an isolated copy of the yuno (see below) with a single call to the command that builds it.

Reproduce it off the live yuno. Copy the store under a scratch root, copy the yuno’s config layers, point environment.work_dir at the scratch root, add "yuno": {"autoplay": true} — nothing plays a yuno without the agent — move the listening ports, and lower io_uring_entries (a second yuno of the same size fails io_uring_queue_init_params() with ENOMEM). Validate the copy before believing a negative: plant a gbmem_malloc(1234) in main() and check the report prints it. A copy that does not leak is then a fact about the difference between it and production, and that difference is the answer.


12. Code pointers

WhatWhere
Severity log APIkernel/c/gobj-c/src/glogger.c
LOG_AUDIT / LOG_MONITORkernel/c/gobj-c/src/glogger.c
Trace emit API (gobj_trace_msg/json/dump)kernel/c/gobj-c/src/glogger.c:778
Global trace level tablekernel/c/gobj-c/src/gobj.c
Per-gclass trace declaration (example)kernel/c/root-linux/src/c_tcp_s.c
Trace mask lookupkernel/c/gobj-c/src/gobj.c (gobj_trace_level)
Per-gobj trace APIkernel/c/gobj-c/src/gobj.c (gobj_set_gobj_trace)
no_trace APIkernel/c/gobj-c/src/gobj.c
Deep tracekernel/c/gobj-c/src/gobj.c
trace_machine printkernel/c/gobj-c/src/glogger.c:1161
FSM dispatch trace siteskernel/c/gobj-c/src/gobj.c
Trace persistence (trace_levels attr)kernel/c/root-linux/src/c_yuno.c
Trace commands exposed by every yunokernel/c/root-linux/src/c_yuno.c
daemon_log_handlers parserkernel/c/root-linux/src/entry_point.c
Log file path builderkernel/c/root-linux/src/yunetas_environment.c
Log line discover() (metadata fields)kernel/c/gobj-c/src/glogger.c:1234
UDP wire formatkernel/c/gobj-c/src/log_udp_handler.c
ievent_gate_stack constantkernel/c/root-linux/src/msg_ievent.h
ievent_gate_stack push/popkernel/c/root-linux/src/msg_ievent.c
logcenter listeneryunos/c/logcenter/src/c_logcenter.c
logcenter commandsyunos/c/logcenter/src/c_logcenter.c
SPA dev-panel rendererkernel/js/gobj-ui/src/yui_dev.js
SPA inter-event callback hookkernel/js/gobj-js/src/c_ievent_cli.js
SPA teardown orderkernel/js/gobj-js/src/c_ievent_cli.js