The async iterator API is quite ergonomic, and allows us to send paths in batches. Sending paths in batches means that we can start processing the paths before we've finished searching, and also makes passing messages between rust and nodejs much faster. I tried quite a few variants, like using a Buffer and manually reconstructing strings, callbacks, etc. The fastest way to pass data between rust and nodejs is by encoding / decoding JSON. It's pretty unintuitive, but this is mainly because rust and C++ (native JSON.parse) are way faster than constructing strings manually with the v8 engine. Both callbacks and iterators are quite fast, and have their own trade-offs. The API for iterators is much more ergonomic however, so that's what I pursued. The performance improvement compared to the current implementation is huge. Up to 2-3x in some cases. ❯ hyperfine 'node example/old.ts' 'node example/stream.ts' 'node example/main.ts' Benchmark 1: node example/old.ts Time (mean ± σ): 2.007 s ± 0.478 s [User: 2.572 s, System: 2.357 s] Range (min … max): 1.619 s … 3.326 s 10 runs Benchmark 2: node example/stream.ts Time (mean ± σ): 1.234 s ± 0.258 s [User: 2.976 s, System: 2.811 s] Range (min … max): 0.913 s … 1.735 s 10 runs Benchmark 3: node example/main.ts Time (mean ± σ): 1.117 s ± 0.267 s [User: 2.425 s, System: 2.484 s] Range (min … max): 0.809 s … 1.680 s 10 runs Summary node example/main.ts ran 1.10 ± 0.35 times faster than node example/stream.ts 1.80 ± 0.61 times faster than node example/old.ts
1.7 KiB
@immich/walkrs
High-performance file tree walker for Node.js, built with Rust and the battle-tested ignore crate from ripgrep.
Background
This project grew out of the need for fast and reliable external library scanning in Immich. Immich needed to scan very large photo libraries efficiently, often containing hundreds of thousands of files across complex directory structures.
By leveraging Rust's performance and the same ignore logic used in ripgrep (one of the fastest file search tools available), walkrs delivers exceptional speed and reliability for file tree traversal.
Installation
pnpm add @immich/walkrs
Usage
import { walk } from '@immich/walkrs';
// Simple usage - walk a directory
const files: string[] = [];
for await (const batch of walk({ paths: ['/path/to/scan'] })) {
files.push(...JSON.parse(batch));
}
// Advanced usage with filtering
const photos: string[] = [];
for await (const batch of walk({
paths: ['/photos', '/backup/photos'],
extensions: ['.jpg', '.png', '.heic', '.webp'],
exclusionPatterns: ['**/.stfolder/**'],
includeHidden: false,
})) {
photos.push(...JSON.parse(batch));
}
Performance
walkrs is designed to handle massive directory trees efficiently. It is greatly affected by multithreading: In benchmarks we have scanned 11M files in under 30 seconds over NFS on a machine with 32 CPU threads available. When restricting walkrs to a single thread, the time for the same task goes up to 208 seconds. Compare this with the single-threaded fast-glob which uses 360 seconds for the same task.