
Peter BednarčíkAlmost every tool that looks at files asks the same question: What's in this directory, how big is...
Almost every tool that looks at files asks the same question: What's in this directory, how big is each file, and when did it last change? Backup tools ask it, build caches ask it, sync clients, indexers, and file watchers ask it, thousands of times a minute.
In Go, the usual answer looks like this:
entries, _ := os.ReadDir(dir)
for _, e := range entries {
info, _ := e.Info() // on macOS and Linux: one more system call per file
use(e.Name(), info.Size(), info.ModTime())
}
On macOS and Linux, that's one call to list the directory, then one lstat for every file. On Windows, the listing already carries sizes and times, but the standard library still does a lot of work per entry. For 1,000 files, it makes 2,000 to 5,000 allocations, depending on the platform.
macOS can return names and metadata in one call, and Linux has cheaper ways to stat a file than using a full path. os.ReadDir with Info uses neither.
So I decided to write a small Go library that does: FlexReadDir. It lists a directory and gives you every file's size and modification time in one pass.
It's part of something bigger I'm working on, but I wanted to show this piece on its own. I also like optimizing software and getting as much performance as I can out of tools when it makes sense.
go get github.com/pbednarcik/flexreaddir/frd
entries, err := frd.ReadDir(".")
if err != nil {
log.Fatal(err)
}
for _, e := range entries {
if e.Type.IsRegular() {
fmt.Printf("%-24s %8d bytes %s\n", e.Name, e.Size, e.ModTime().Format("2006-01-02 15:04"))
}
}
ReadDir returns entries sorted by name, exactly like os.ReadDir. Each Entry is a plain struct. Size, times, and file id are filled for regular files only; for directories, links, and other types they are zero:
type Entry struct {
Name string
Type fs.FileMode
Size int64
ModTimeNanos int64
ChangeTimeNanos int64
FileID uint64 // inode / file id; 0 if unknown
}
Because Entry is comparable, comparing a file's recorded metadata with the last scan is a single ==:
for _, e := range entries {
if e.Type.IsRegular() && previous[e.Name] != e {
fmt.Println("changed:", e.Name)
}
}
With both the change time and the file id in there, a scanner can even tell a file that was rewritten in place from one that was replaced by another file with the same name, where the filesystem gives a persistent file id.
| Platform | Files |
ReadDir + Info
|
Readdir |
FlexReadDir | Faster |
|---|---|---|---|---|---|
| macOS | 1,000 | 1,382 µs | 947 µs | 362 µs | 3.8× |
| Linux (VM) | 1,000 | 603 µs | 465 µs | 376 µs | 1.6× |
| Windows | 1,000 | 164 µs | 161 µs | 131 µs | 1.3× |
FlexReadDir here is AppendDir, unsorted and reusing its result slice, as a walker would use it. "Faster" is against os.File.ReadDir with Info for each entry; os.File.Readdir is the older API that returns a FileInfo per entry. Every row is the median of ten interleaved runs on a warm file cache, so it shows warm-cache speed, not disk speed. The Mac is an M5 Max on APFS, the Windows box a Ryzen 9 9950X on NTFS, and Linux runs in a VM on that same Mac, so treat its row as a direction, not a size.
The gain is different on each platform, because the standard library does more extra work on some systems than on others. The macOS result surprised me the most. getattrlistbulk is a real bulk call, and it shows.
On macOS, it uses getattrlistbulk. You tell the kernel which attributes you want (name, type, size, times, file ID), and it returns records for many files at once. os.ReadDir + Info calls lstat for every file, so most of the 3.8× comes from removing those calls.
Linux has no bulk stat, so FlexReadDir still calls statx once per regular file. The savings come from everything around it. The stat is relative to the open directory instead of a full path; it asks only for the fields it needs, and it reads the getdents64 buffer directly without allocating per entry.
I plan to run it in Docker, so I made sure it works in containers. Some seccomp profiles block statx with EPERM. In that case, FlexReadDir falls back to fstatat, and there is a test with a real seccomp filter for it.
Windows already returns sizes and times with the listing (FILE_ID_EXTD_DIR_INFO), so there is little to gain. The savings come from less work per entry in user space and faster file-name decoding.
No cgo on any platform.
I wrote my first small SIMD programs at university. One of them, for a course assignment, calculated a matrix determinant with AVX, and I compared it with a plain C version, so I saw how fast SIMD can be.
Windows returns file names in UTF-16, and Go strings are UTF-8, so every name must be converted. Most names are plain ASCII, which fits SIMD well. It checks 16 or 32 characters at once and, if they are all ASCII, packs them in one step.
The decoder has three versions: AVX2, SSE2, and plain Go. The right one is picked at runtime, and -tags purego forces plain Go. On long names, like in WinSxS, where names average 94 characters, AVX2 decodes about 4× faster than plain Go, and listing that directory is 17 to 20% faster. All three are tested exhaustively and fuzzed against each other.
This stage probably took the most time. Constant reruns, long Google searches, documentation checks, and a lot of waiting on the rebuild-and-retest loop. But that's the price of optimizing.
After the obvious wins, I wanted to know what was left, so I profiled listings down into the kernel. On Windows, I used the Windows Performance Recorder with Microsoft's public symbols.
In my profiles, FlexReadDir's own code was about 4% of the time on macOS, 5% on Linux (a tree walk in Docker), and under 13% on Windows. The rest is mostly the kernel.
Two other things stood out: how much work NTFS does on its own, and how quick Linux is with file name lookups. The Linux part didn't really surprise me. I remember Linus Torvalds saying in a talk that he spent a lot of time keeping cache misses in file name lookups as low as possible.
On a warm cache, 40 to 49% of a Windows listing is NTFS reading ahead in the directory index, and I couldn't do anything about it. Or I just haven't figured it out yet. Access hints when opening the directory didn't change anything. A 16-times-bigger buffer (1 MiB instead of 64 KiB) wasn't faster either, and on the system drive (C:) it was 4.8% slower. In the end, I left it to Windows. At some point, you have to call it good enough, and I'd reached it.
I come from the C# world, but I've always been interested in memory and being close to the metal.
Listing 1,000 files takes 4 or 5 allocations when you reuse the result slice. The names from one read share one allocation, and a Lister keeps its 64 KiB buffer, so when you list many directories you can reuse both:
var l frd.Lister
var entries []frd.Entry
for _, dir := range dirs {
f, err := os.Open(dir)
if err != nil {
return err
}
entries, err = l.AppendDir(entries[:0], f)
_ = f.Close()
if err != nil {
return err
}
// use entries before the next iteration overwrites them
}
Because names share memory, keeping a Name for a long time keeps the others alive too. Use strings.Clone if that matters.
It lists one directory and doesn't walk trees. If you need a walk, an example on pkg.go.dev builds one on top of it.
I don't have Linux on real hardware, an Intel CPU, or NFS and SMB shares to test on. If you have any of them, clone the repository and run this. It takes about a minute:
for i in $(seq 10); do go test ./frd -run '^$' -bench 'AppendDir' -benchmem -count 1; done
On Windows, in PowerShell:
1..10 | ForEach-Object { go test ./frd -run '^$' -bench 'AppendDir' -benchmem -count 1 }
The benchmark creates its files in the temp directory, so to test an NFS or SMB share, point TMPDIR (TMP on Windows) at a folder on the share first. Open an issue with the output and a line about the machine.
It requires Go 1.27 or later, runs on macOS, Linux, and Windows, and is MIT-licensed.
Go was an experiment for me, a way to try something new. It has a garbage collector like C#, so it felt familiar. In the end, it was a great way to learn Go, and I'm sticking with it, at least for now.
Let me know what you think, especially if you write Go. The language is still new to me, and I tried really hard not to write C# in Go. Hopefully I managed :D If something in the API or the code looks odd, tell me.
Thanks for reading.
I used AI to help edit this article.