The following is part of a series of posts about 2026 summer intern projects—for more, see “What the interns have wrought, special jumbo 2026 edition”
At Jane Street, we have many data processing workloads that require high-throughput access to large amounts of storage. Historically, we’ve been able to rely on off-the-shelf products provided by external vendors (e.g. Dell Isilons or VAST Storage) to service this need, but dataset sizes have grown and research benefits from ever faster access to more storage, causing us to hit scaling limits in these products. So, we developed our own in-house storage system called Depot, which aims to provide a highly scalable object store (think S3, optimized for our use cases) to service research workloads.
As is our custom, Depot is written in OCaml, and the primary client library for accessing Depot is in OCaml as well. However, various components in our core trading research infrastructure expect to access files using a regular POSIX filesystem interface. Some can be migrated to use Depot without much difficulty, but others would require deep architectural changes. To ease onboarding these tools to Depot, we decided to create a Filesystem in Userspace (FUSE) server to expose the Depot object store in a way that supports the POSIX interface. Since our primary use case is servicing research workloads, the FUSE server needs to be high-performance, but can be read-only.
During his Summer 2026 internship, Linux Engineer Kian Kasad developed an OCaml-native FUSE library and used it to build a high-performance read-only FUSE server that exposes the Depot filesystem.
FUSE in OCaml
For our purposes, FUSE is a protocol that allows the Linux kernel and a regular userspace program to work together to expose a POSIX filesystem. Once some initial setup and handshaking is done, it works by having the kernel translate requests performed against the filesystem to messages for the userspace program, which replies to the kernel, which then delivers the final result to the program that performed the IO.
FUSE servers typically use libfuse, the C reference library for interacting with FUSE on
Linux. However, libfuse is challenging to integrate neatly with our internal OCaml
ecosystem because it performs blocking IO, which clashes with our
Async runtime. It
also requires an FFI, which is annoying, error-prone and can have performance
implications. Instead, we decided to implement the FUSE protocol ourselves. The resulting
library can then be used for Depot, as well as many other internal projects that also use
FUSE.
Unfortunately, the FUSE protocol has relatively little documentation, official or unofficial. Kian spent a lot of time digging into kernel code for each message type to ensure that he understood the semantics of the message or even whether that message applies to FUSE (as opposed to FUSEBLK or CUSE or Virtio-FS). He also wrote comprehensive tests to ensure that our library was serialising and deserialising messages correctly. For this, instead of trying to theoretically enumerate the edge cases we might encounter, he just altered an existing libfuse-based implementation to capture the requests and responses it saw, manually ran a bunch of operations against this FUSE server to capture as many message types as possible, and then used these as test data for our serializers and deserializers to confirm that they behaved correctly. (This is a common pattern at Jane Street, where you capture production data for realistic end-to-end tests; some systems even build this in so that you can retroactively mint such a test, as when there’s a tricky bug in prod.)
Once we got the protocol working, we started actually talking to the kernel. FUSE servers
begin this process by opening /dev/fuse. This is a character device, which means its
driver is free to implement the file operations API however it wants. And we learned the
hard way that the /dev/fuse device requires reading / writing exactly one FUSE message
in each read() / write() syscall (there’s no official docs about this!). But since
other common ways of doing IO, like regular files, TCP sockets, and pipes don’t behave
this way, our standard libraries assume short reads and partial writes are a fact of life
and build upon these to do internal buffering and transparent retries. So we had to comb
through our libraries to find functions that we could guarantee matched the behaviour of
the FUSE device; this involved reading the source all the way down to the C stubs to
ensure they were only issuing one read() / write() syscall. As a fun example, here’s
how we write the messages:
Fd.syscall_in_thread t ~name:"fuse-write" (fun fd ->
(* Even though we require the full message to be written, we don't set [~min_len],
because that causes [Bigstring_unix.write] to issue multiple writes. *)
Bigstring_unix.write fd buf ~pos ~len:msg_len)
Turns out that if you pass a non-zero ~min_len to Bigstring_unix.write, it’s liable to
loop and write() again if the first write was partial for some reason, so we had to be
very careful!
Depot via FUSE
Next, we implemented a FUSE server for Depot. This is a tricky problem—we have to translate the semantics of Depot, an object store system with immutable objects and a very different approach to handling filesystem metadata, into the POSIX semantics that Linux wants. This is fundamentally a policy decision, so we worked closely with the developers of the Depot storage system and incorporated how researchers use the filesystem to try to develop a mapping that preserves the useful semantics of POSIX filesystems while also faithfully expressing the ways that Depot is fundamentally different from a POSIX filesystem.
A simple example was the length of file/directory names. Depot supports names up to 7,000 characters, but FUSE supports only up to 1,024, and much UNIX tooling supports only 256. The naïve approach would be to return errors for these, but this would cause friction when working with existing Depot datasets using FUSE. To solve this, we map long entry names to shorter UNIX-compatible names when browsing directories, then map the short names back to their long equivalents when users operate on these paths.
Another problem is the way Depot handles paths. POSIX filesystems typically support hardlinks, which allow the same underlying file to live in two places. However, POSIX does not allow hardlinking directories, as it can break several invariants by allowing cycles or making parent paths ambiguous. In Depot on the other hand, “bundles” (similar to directories) relax this restriction and are allowed to be referenced by multiple paths to simplify scaling. To expose this functionality in POSIX, we felt the best tradeoff was to pretend as if each view of the bundle was a different filesystem object. This was implemented by decoupling the inodes returned by the FUSE from the underlying bundles, and instead keying them just based on the path used to access the bundle.
Performance optimization
Once the Depot-backed FUSE was operational, benchmarks showed that it was only able to serve ~30 MiB/s when reading large files. This would be a stark downgrade from the performance of our current storage systems.
One source of slowness came from having to serialize multiple reads to the same file, since the Depot client API’s documentation warns against calling a reader function in parallel. However, after digging into the implementation, we realised that it would handle parallel operations just fine. We confirmed with Depot devs that we can actually perform multiple reads in parallel, which then doubled the observed throughput since it allowed the Linux kernel’s built-in readahead logic to work properly.
The most impactful change, though, was reducing unnecessary copying of data. We had
initially represented data using OCaml’s string type, which meant a read operation would
copy the data twice: once to hydrate a string from an IO buffer, and once to go back to an
IO buffer. When reading a 4 GiB file, this would result in ~8 GiB of memory allocation. It
turned out that 92% of the program’s overall allocations came from such operations. To
save memory, we started passing buffers around directly, allowing data to be moved without
copying. This required reasoning about the lifetime of buffers to ensure no buffer was in
use by two operations at once. But the bookkeeping was worth it: after these changes, we
reached read speeds of 1.3 GiB/s. We didn’t hit a CPU bottleneck or saturate the NIC, so,
picking up where Kian left off, we think the next limitation is likely Linux’s small
readahead window size.