Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

fs_watcher

inotify-based filesystem watcher used by timeranger2 and by the C_FS gclass to react to on-disk changes.

Source code:

fs_create_watcher_event()

fs_create_watcher_event() initializes a new file system watcher event, monitoring the specified path for changes based on the given fs_flag. The event is associated with the provided yev_loop and invokes the specified callback when triggered.

fs_event_t *fs_create_watcher_event(
    yev_loop_h     yev_loop,
    const char     *path,
    fs_flag_t      fs_flag,
    fs_callback_t  callback,
    hgobj          gobj,
    void           *user_data,
    void           *user_data2
);

Parameters

KeyTypeDescription
yev_loopyev_loop_hThe event loop handle in which the watcher event will be registered.
pathconst char *The directory or file path to monitor for changes.
fs_flagfs_flag_tFlags specifying monitoring options, such as recursive watching or file modification tracking.
callbackfs_callback_tThe function to be called when a file system event occurs.
gobjhgobjA generic object handle associated with the watcher event.
user_datavoid *User-defined data passed to the callback function.
user_data2void *Additional user-defined data passed to the callback function.

With FS_FLAG_MOVED_AS_DELETED a subdirectory renamed away (out of the watched directory, or to another name in it) is told as FS_SUBDIR_DELETED_TYPE, by its OLD name -- for a watch that is not recursive (in a recursive one the moved directory’s own watch would keep its old path). The master of a timeranger2 topic watches disks/ this way: a reader closes its feed by renaming disks/<rt_id>/ away before removing it.

fs_event_t *fs_event = fs_create_watcher_event(
    yev_loop, "/yuneta/store/db/topic/disks", FS_FLAG_MOVED_AS_DELETED,
    master_fs_callback, gobj, tranger, NULL
);
// mv disks/rtA disks/.closing.123-4567.1  ->  FS_SUBDIR_DELETED_TYPE, filename "rtA"

The descriptors of FS_FLAG_DIR_FDS live as long as their watch -- a follower holds one per key directory of each feed -- so the watcher says when the ones held by every watcher of the process reach half of the soft open-files limit (“Directories watched through descriptors: half of the open-files limit”, with dir_fds of the process, dir_fds_of_this_watcher and soft_limit), and says it again only after they fell under the half. The limit is the process’s: up to 7.25.21 each watcher counted only its own, so four followers of 400 directories each, under a limit of 1024, never said it. A directory whose descriptor cannot be opened (EMFILE) is watched by its path; the first failure is an ERROR, the next ones are only counted, and the count is said when a descriptor opens again (“Directories watched through their descriptor again”, watched_by_path). A yuno raises its own soft limit to its hard one at its start (C_YUNO limit_open_files, 0 by default).

With FS_FLAG_RECURSIVE_PATHS every subdirectory is watched, hidden ones (.name) included, whenever it appears: there at the start, made later (IN_CREATE), met by the pass after an overflow, or by the re-watch of one that could not be watched. Up to 7.25.22 the first walk left the hidden ones out while the others took them, so a .cache there before the watch was deaf, and one made after it was heard.

mkdir("/data/.h", 0770);            // before the watch: watched too
fs_event_t *fs = fs_create_watcher_event(loop, "/data", FS_FLAG_RECURSIVE_PATHS, cb, gobj, 0, 0);
fs_start_watcher_event(fs);         // a file made in /data/.h is heard

With FS_FLAG_DIR_FDS each SUBDIRECTORY watched is opened first and watched through that descriptor, so the watch and the descriptor are one inode: see fs_watcher_dir_fd(). Every event carries event_wd (the watch of directory) and, for FS_SUBDIR_CREATED_TYPE, subdir_wd (the watch just set on the directory created, -1 if it was gone).

Returns

Returns a pointer to a newly allocated fs_event_t structure representing the watcher event, or NULL on failure -- a path that is not a directory, no inotify instance, or a root that cannot be watched (ENOSPC at fs.inotify.max_user_watches, logged). Up to 7.25.20 a root that could not be watched gave a watcher all the same, running and watching nothing.

With FS_FLAG_BATCH_END the owner is also called with FS_BATCH_END_TYPE after each batch read from inotify (offset = where the batch ends), and after each slice of the pass that follows an overflow (offset = where the stream is), and every event carries offset and offset_end, its place in the watcher’s stream: what the owner left for “when the stream is past here” can be done there. A timeranger2 follower defers the scan of a key directory that way. It notes the directory at its event, and asks where the queue ends once, at the end of the batch: fs_queued_events_end() walks the whole inotify queue, and asked at each new directory it made a flood of 69632 new keys quadratic (18 s of a drain of 21). The order is the contract: first LOOK at each directory (who it is: the watch it was seen with, and its descriptor -- see FS_FLAG_DIR_FDS; without one, inode and birth), THEN ask where the queue ends, and read a directory only if it is still the one looked at. A change made by another process after the look is either queued before the answer (and read before the directory is) or it changed the directory (and the read is skipped). With the question first and the look after, a directory removed and made again between the two is read as the new one, while the event of its removal is queued past the answer. The window is not small: the owner’s own callbacks for the directories before it in the batch run there.

A file created in a directory that is noted or placed is left to the read of that directory, which takes all its files in order. Read at its own event, a second file of a new key came before the first.

fs_event_t *fs = fs_create_watcher_event(
    yev_loop, path, FS_FLAG_RECURSIVE_PATHS|FS_FLAG_BATCH_END, my_fs_callback, gobj, NULL, NULL
);
...
case FS_SUBDIR_CREATED_TYPE:
case FS_RESCAN_DIR_TYPE:
    note_the_directory(my, fs_event);               // cheap: nothing asked here
    break;
case FS_FILE_CREATED_TYPE:
    if(directory_is_noted_or_placed(my, fs_event)) {
        break;                                      // its read takes this file, in order
    }
    read_the_file(my, fs_event);
    break;
case FS_BATCH_END_TYPE:
    read_the_placed_directories_due(my, fs_event->offset);  // each one if still the one looked at
    if(has_notes(my)) {
        look_at_the_noted_directories(my);          // FIRST: each one by its watch and descriptor
        uint64_t until = fs_queued_events_end(fs_event);    // THEN the question, once
        place_the_notes(my, until);                 // read each when the stream is past `until`
    }
    break;

Notes

The created watcher event must be started using fs_start_watcher_event() to begin monitoring. When no longer needed, it must be stopped using fs_stop_watcher_event(), which will also free the associated resources.


fs_start_watcher_event()

fs_start_watcher_event() starts monitoring the specified file system event. This enables notifications for file and directory changes.

int fs_start_watcher_event(
    fs_event_t *fs_event
);

Parameters

KeyTypeDescription
fs_eventfs_event_t *Pointer to the file system event structure to be started.

Returns

Returns 0 on success, or a negative error code on failure.

Notes

Once started, the event will trigger the associated callback when file system changes occur. Use fs_stop_watcher_event() to stop monitoring and release resources.


fs_stop_watcher_event()

fs_stop_watcher_event() stops the given file system watcher event and destroys the associated fs_event_t instance.

int fs_stop_watcher_event(
    fs_event_t *fs_event
);

Parameters

KeyTypeDescription
fs_eventfs_event_t *Pointer to the fs_event_t instance representing the file system watcher event to be stopped and destroyed.

Returns

Returns 0 on success, or a negative error code on failure.

Notes

Once fs_stop_watcher_event() is called, the fs_event_t instance is destroyed and must not be used again.


fs_queued_events_end()

fs_queued_events_end() says where the events the kernel holds for the watcher BY NOW end, in the watcher’s stream of events. Every event the watcher hands over carries its own place in that stream, fs_event->offset (the bytes read from inotify before it). An event handed over later with an offset below the answer was already queued when the question was asked: what it says may be what the owner has just read from the disk.

uint64_t fs_queued_events_end(
    fs_event_t *fs_event
);

Parameters

KeyTypeDescription
fs_eventfs_event_t *The watcher.

Returns

The offset in the stream where the events queued by now end: the rest of the batch being walked, if one is -- or else what a read the kernel completed and the loop has not delivered took --, plus what the kernel still holds (FIONREAD). 0 for a NULL watcher. If FIONREAD fails the error is logged and the answer is past everything the kernel can hold for the fd (max_queued_events events of the largest size, from /proc/sys/fs/inotify/max_queued_events), plus a read whole outside a batch: never short. Up to 7.25.21 only a read whole was added (inside a batch, at first, nothing), and with more than a read queued the owner did too soon what it left for after those events. If FIONREAD fails again when the watcher comes to close that end, the watcher is gone (see below).

Notes

Exact, asked from the owner’s own callback and of another watcher alike. Between two batches the kernel completes the watcher’s read at any return to user space (an interrupt’s too): its events are out of the kernel’s queue and not yet handed over. The completion is looked for in the loop’s ring (yev_get_waiting_completion()) before and after asking the kernel; the same answer both times means nothing moved in between (there is one read at a time, and once completed it waits for the loop). Only when completions overflowed the ring, and whether one of this read waits cannot be seen, is a read counted whole: the answer is then past the end, never short.

An answer past the end is closed by the watcher itself. It notes it, and on the next turn of the loop (and after each batch while it is open) it looks again: once no completion of its read waits in the ring and the kernel holds nothing, everything queued when it was asked has been handed over, so the stream JUMPS to that answer and the owner is called with FS_BATCH_END_TYPE at it (with FS_FLAG_BATCH_END). The offsets after it continue from there: an offset is a place in the stream to compare, not a count of bytes read. Up to 7.25.21 nothing closed it, and an owner waiting for “the stream past here” waited for about 8 KB of unrelated events -- on a quiet watcher, for ever.

case FS_BATCH_END_TYPE:
    if(owner->scan_pending && fs_event->offset >= owner->scan_until) {
        scan_the_directory(owner);  // reached, even if no event came after it
    }
    break;

It costs what the queue holds: FIONREAD walks the whole inotify queue. Ask it once per batch (FS_BATCH_END_TYPE), not once per event: asked at each of 69632 new directories it was 18 s of a drain of 21. A read takes up to 32 events of the longest name (up to 7.25.20 one: a backlog of 65536 events was ~8000 batches, and as many questions at their ends).

A timeranger2 follower uses it to tell apart what it already said from what is new. At an overflow it reads keys/ and tells the keys gone from there deleted; the signal of such a delete can still be in the queue, behind the overflow, and must not be told again:

case FS_OVERFLOW_TYPE:
    list_what_is_on_disk_and_tell_it(owner);       // what the lost events said
    owner->told_until = fs_queued_events_end(fs_event);
    break;

case FS_SUBDIR_DELETED_TYPE:
    if(fs_event->offset < owner->told_until && already_told(owner, fs_event->filename)) {
        break;  // queued before the owner read the disk: said already
    }
    tell_deleted(owner, fs_event->filename);
    break;

fs_watcher_dir_fd()

With FS_FLAG_DIR_FDS, the descriptor of the directory watched under wd (an event’s event_wd or subdir_wd). It is the very inode of that watch: the directory is opened BEFORE it is watched, and watched through the descriptor (/proc/self/fd/N), so a directory removed and made again under the same path -- which ext4 gives the same inode number, and on 6.x kernels the same birth time -- is never taken for it. openat() / unlinkat() through it cannot reach another directory: in a directory gone they fail with ENOENT.

int fs_watcher_dir_fd(
    fs_event_t *fs_event,
    int wd
);

Parameters

KeyTypeDescription
fs_eventfs_event_t *The watcher.
wdintA watch: fs_event->event_wd, fs_event->subdir_wd.

Returns

The descriptor, owned by the watcher: do not close it, and take it again at each event (it may be closed at the next one). -1 and errno:

Notes

A descriptor open on a directory holds its inode: the kernel then holds back its IN_DELETE_SELF and IN_IGNORED until the descriptor is closed. The watcher closes it when the directory’s parent reports it deleted (IN_DELETE|IN_ISDIR) and it has no link left (st_nlink 0), and when a new directory is watched at its path; then its events come, and the watch goes as without the flag. The cost is one descriptor per subdirectory watched (a timeranger2 follower: one per key directory of each feed), and an open() per directory watched.

A timeranger2 follower takes the link of a new record through the descriptor of the directory the event came from, so an event of a key directory deleted and made again before it was read reaches nothing:

case FS_FILE_CREATED_TYPE: {
    int dir_fd = fs_watcher_dir_fd(fs_event, fs_event->event_wd);
    if(dir_fd >= 0) {
        int fd = openat(dir_fd, fs_event->filename, O_RDONLY|O_CLOEXEC|O_NOFOLLOW);
        if(fd < 0 && errno == ENOENT) {
            break;      // consumed, or its directory is gone: not this one's
        }
        unlinkat(dir_fd, fs_event->filename, 0);
        read_the_life_of(fd);
        close(fd);
    } else if(errno == ENOTSUP) {
        read_by_path(fs_event->directory, fs_event->filename);
    }
    // ENOENT: the directory is gone, its delete comes
    break;
}

Queue overflow (IN_Q_OVERFLOW)

Each watcher owns one inotify instance with a bounded kernel event queue (fs.inotify.max_queued_events; the deb/rpm packagers set 65536). Under a burst the kernel drops the events that do not fit and signals a single IN_Q_OVERFLOW: from there the watcher cannot know every change it missed.

What the lost events said is still on the filesystem, so the watcher recovers in place (since 7.25.9; in slices since 7.25.10):

  1. a WARNING, “inotify IN_Q_OVERFLOW: events lost, rescanning the watched tree”, with the watched path;

  2. the owner’s callback is called ONCE with FS_OVERFLOW_TYPE (directory = the watched path): do there what is global and cheap. An owner may stop the watcher there (fs_stop_watcher_event()): then no pass is started, and the watcher goes when the batch ends (up to 7.25.20 the pass was still set up -- its index built, its timer created and armed -- and thrown away);

  3. a pass over the tree follows: every directory, the root included, is handed to the owner as FS_RESCAN_DIR_TYPE (directory = that directory) -- read it again, what it holds may never have been told. In a recursive watch a directory born while its IN_CREATE was dropped is watched before it is handed over, and so is one deleted and created again meanwhile: every directory of the pass is watched again (inotify_add_watch() on an inode already watched returns its wd), and a wd that differs from the one the table holds for that path replaces it. Up to 7.25.20 a path found in the table was taken as watched, and a directory reborn during an overflow (another inode, its IN_IGNORED lost) was never heard again; the table also let an IN_IGNORED pass without taking out its wd, which it does now. The ROOT is watched again too, in a watch that recurses and in one that does not: up to 7.25.20 the pass watched again the directories it met, never the root it started from, and a root deleted and created again during an overflow went deaf. At the end of the pass a directory of the table that the pass did not meet and that is no longer there is stopped too. The entry of a wd stopped goes with its IN_IGNORED; when that was dropped with the overflow, it goes once the stream is past where the IN_IGNORED would have come (fs_queued_events_end() at the stop): up to 7.25.21 it stayed for good. The watcher does not follow moves: a directory that is not there is gone for it;

  4. the pass runs a slice of 20 ms per loop turn, and an INFO closes it: “watched tree rescanned after lost inotify events”, with directories, ms, and where that time went: slices, ms_owner (in the owner’s callbacks), ms_watcher (the walk itself), ms_loop (the loop’s own work between slices) and max_loop_ms (its longest turn). Another overflow during a pass schedules one more whole pass after it (starting again would starve the end of the tree under overflows that keep coming).

In 7.25.9 the owner rescanned the whole tree inside the one FS_OVERFLOW_TYPE call. A timeranger2 follower of 50000 keys on the busy disk of a central took 56 s, then 243 s -- the yuno deaf to its agent, its commands and its timers all that time. The slices keep it answering; the pass costs the same.

What the pass costs is mostly the owner’s: the watcher’s own part is a readdir() per directory and a lookup in an index of the watched paths, built once per pass. 7.25.10 rebuilt that index in every slice -- 50000 paths every 20 ms -- so its own cost grew with the tree, and on the central a pass over 50501 directories took 7 minutes. tests/c/timeranger2/test_fs_watcher_overflow measures it (the pass less the owner’s time, per directory) on a tree of 4096 directories and on one of 69632, on the same machine, and fails when the big one costs more than twice the small one plus 20 us: with the index once per pass, 14 and 20 us here; rebuilt per slice, 31 and 437. (Up to 7.25.22 it failed above a fixed 200 us, which a slower machine reached with the code right.)

A pass closes with its own account (since 7.25.12), so a long one says why:

"msg": "watched tree rescanned after lost inotify events",
"directories": 69633, "ms": 14374, "slices": 612,
"ms_owner": 11169, "ms_watcher": 887, "ms_loop": 2318, "max_loop_ms": 17

Here the owner took most of it (a test’s owner that sleeps 100 us per directory); a large ms_loop would say the slices were waiting for a busy loop instead.

What it said on yunovatios’ central, a timeranger2 follower of 50501 keys while its master wrote a backlog of ~9 GB (2026-09-27): passes of 142-335 s, 65-78 % in the owner, 22-35 % in the loop, and 1.1-1.7 s in the watcher -- whatever the length of the pass. The owner’s time is the records it hands over: the same tree with nothing pending took 1 s, 0.35 s of it in the owner. So a long pass is a follower working through a backlog at its own pace (there, ~80 % of a core), not a slow walk; it ends about a minute after the master stops writing, and the loop never waited more than 239 ms for a slice.

Every owner handles both -- in its callback, before anything that reads the type as bits:

PRIVATE int my_fs_callback(fs_event_t *fs_event)
{
    switch(fs_event->fs_type) {
        case FS_OVERFLOW_TYPE:
            // events were lost: what is global and cheap (a pass follows)
            break;
        case FS_RESCAN_DIR_TYPE:
            // one directory of the tree: list it again
            rescan_my_dir(fs_event->gobj, (const char *)fs_event->directory);
            break;
        case FS_FILE_CREATED_TYPE:
            ...
    }
    return 0;
}

What the owners of the tree do:

A subdirectory that cannot be watched (ENOSPC)

A subdirectory met by a recursive watch whose watch cannot be made -- out of watches (fs.inotify.max_user_watches, ENOSPC) or of kernel memory (ENOMEM) -- is handed as created with subdir_wd -1, and nothing made in it is heard. It is kept, and tried again at the end of each batch of the watcher (64 per batch); once its watch is made it is handed AGAIN as created, now with its subdir_wd, at the batch’s end, so its owner reads what was made in it meanwhile. One gone by then is forgotten (its parent said it). Under FS_FLAG_RECURSIVE_PATHS what was made under it was not heard either: its subdirectories not watched are queued behind it and tried in the same way, parent first, so each one is handed as created after its parent. Every try counts against the batch’s 64, and an ENOSPC/ENOMEM ends the batch’s tries: a large subtree is watched over several batches. What is not bounded is the listing of one directory watched again: it is read whole in that turn (a keys/ of 100 000 keys is one readdir of it). A directory gone between its watch and its listing says nothing (its parent’s IN_DELETE does). The first failure is an ERROR, “Cannot watch a directory, out of inotify watches or memory: tried again at each batch (and the next ones that fail, counted)”; at the end of the batch where none is left, a warning, “Directories watched again: every one that could not be is watched now”. Up to 7.25.22 a subdirectory made meanwhile was never watched. Up to 7.25.21 it was an ERROR per directory, and the directory was never watched: a timeranger2 follower took a key directory for gone, and lost every record of the key.

case FS_SUBDIR_CREATED_TYPE:
    if(fs_event->subdir_wd < 0) {
        break;  // gone, or not watched yet: if it lives, it comes again with its watch
    }
    read_what_is_in(fs_event->directory, fs_event->filename);
    break;

The retry needs a batch: a watcher whose only activity is inside the directory it cannot watch hears nothing to retry on (nothing made there reaches it), which the ERROR says. Raise fs.inotify.max_user_watches.

When the watcher goes (FS_WATCHER_GONE_TYPE)

A watcher whose read FAILS, cannot be armed again, or is canceled by another than its owner, is over: an ERROR, “inotify read FAILED: the watcher is gone”, “inotify read cannot be armed again: the watcher is gone”, or “inotify read canceled, and not by its owner: the watcher is gone”, with path (and errno of a read), then the owner’s callback is called once with FS_WATCHER_GONE_TYPE (directory = the watched path), and the watcher is destroyed when the call returns. Nothing else comes. So is a watcher that cannot count the kernel’s queue (FIONREAD failing) when it comes to close an end said past the stream: “The events queued by the kernel cannot be counted: an end said past the stream cannot be closed, the watcher is gone” -- up to 7.25.21 that end was left open, and on a quiet watcher never reached. The owner drops every pointer it keeps to it -- and must not stop it: it is freed. An owner that stopped the watcher itself (fs_stop_watcher_event()) is not told. Up to 7.25.20 the watcher went silently (the failure logged only under a trace), and its owner kept a pointer to freed memory: a timeranger2 feed stopped it again when closed.

No shutdown cancels a watcher behind its owner. yev_loop_stop() cancels every operation of the loop, but a yuno calls it after its loop ended (yuno_shutdown() only resets it), its services stopped and its gobjs ended -- every watcher stopped by its owner, C_FS in mt_stop, the trangers of the services in theirs -- and destroys the loop without running it (entry_point.c). The tools either start their tranger without a loop (no watcher), or shut it down before yev_loop_stop() (tr2list and treedb_list --follow), or run no loop after it (tr2migrate). And a loop run after yev_loop_stop() delivers nothing behind the stop’s own completion: it breaks there and leaves it at the head of the ring. So the message means a cancel from outside fs_watcher -- an order broken -- and the owner is still told (tests/c/timeranger2/test_rt_disk_watcher_gone makes one on purpose, with yev_stop_event() on the watcher’s read).

case FS_WATCHER_GONE_TYPE:
    priv->fs_watcher = NULL;    // freed when this returns
    gobj_log_error(gobj, 0,
        "function",     "%s", __FUNCTION__,
        "msgset",       "%s", MSGSET_SYSTEM,
        "msg",          "%s", "the watch is gone",
        "path",         "%s", fs_event->path,
        NULL
    );
    break;

What the owners do: a timeranger2 rt_disk feed drops its watcher and its accounts of deletes, and says it is deaf (“rt_disk feed deaf: its watcher is gone”); every reader of the feed’s watcher skips a feed without one. The master’s watch of disks/ says new feeds are not heard any more. C_FS says the path is not watched any more (its size_dl_watch reads 0). utils/c/fs_watcher prints “Watcher gone”.

Up to 7.25.8 an overflow aborted the yuno, to be relaunched and reload clean. Under a sustained burst the reload met the next overflow: in yunovatios’ stress test of its central, a db_history_ce following a db_tracks_ce at ~3000 records/s over ~60000 keys aborted five times in two hours, each relaunch spending 2-3 minutes catching up before falling again.

A follower also hears its OWN consumption: each link it removes is an IN_DELETE in a watched key directory, which it ignores. Half of what fills its queue is that echo, which is why a single burst can overflow it twice.

tests/c/timeranger2: