From jameshan.net

Resilient Infrastructure at the Frontier of BCI

I’m interning at Neuralink, building 0-to-1 infrastructure on a small team that works on everything from platform engineering and hermetic builds to distributed training infrastructure. On this page I want to share the problems a team like ours has to think through when building infrastructure from scratch, and how the rest of the field approaches them.

Infrastructure is the heartbeat of an engineering organization. Every change an engineer merges, every model a researcher trains, and every service a team ships passes through it, and its rhythm sets the rhythm of everything else: how often code reaches production, how long an experiment waits for compute, how quickly a mistake can be undone. AI makes this more true, not less. When agents can write code in minutes, the bottleneck moves to everything that happens after the code is written: building it, testing it, running it, and knowing that it works.

Where it runs

The first decision is whose machines. The cloud charges a premium for the ability to change your mind about capacity: you can double your servers tonight and give them back tomorrow. With spiky load and no one to run hardware, that option is worth a lot. With steady load, sensitive data, and racks you already own, it buys much less.

37signals is the best-documented case. After leaving the cloud, they cut their annual bill from $3.2 million to $1.3 million and now project more than $10 million saved over five years.⁠Their own caveats: they already had data-center space and power, and a dedicated operations team. A company starting from nothing would pay for both. Cloud providers now waive exit fees for customers leaving entirely, but everyday data transfer out is still charged. What actually traps people is data gravity: storing data is cheap, moving it out is not.

Regulation is often cited as a reason to stay on your own hardware, but regulators don’t ask for that. They ask for evidence: exactly what software ran, how it changed, and why.⁠In medical devices, the FDA’s 2025 guidance on Computer Software Assurance covers internal tools used in production, which includes a deploy pipeline, and asks for validation proportional to risk rather than blanket documentation. Owning the whole pipeline makes that evidence easier to produce. For a company like Neuralink, where the work is proprietary and the data is sensitive, keeping most things on its own machines is the natural default, and the rest of this page is shaped by that.

What makes a build trustworthy

The build is where code becomes something you can run, and the question is whether you can trust what comes out. A build is hermetic when its output depends only on inputs you declared: pinned compilers and libraries, a sandbox, and no network access. It is reproducible when someone else can rebuild it from the same source and get the same bytes. Hermeticity is the method, and reproducibility is the proof, because it is the only way a third party can check a binary instead of trusting whoever built it.⁠Debian reports that over 95% of its packages for its current release are reproducible when rebuilt and compared against the official binaries. The usual culprits for mismatches are timestamps, file ordering, and build paths embedded in the output.

Two systems take this seriously. Nix names every build output by a hash of everything that went into it. Bazel, which grew out of Google’s internal build system, turns a codebase into a graph of hashed actions that can be cached, shared, and run on a remote cluster. Both are powerful, and both pay back roughly in proportion to the team running them. A 2024 study of GitHub projects found that 11% of those that adopted Bazel later abandoned it, after a median of 638 days.⁠Bazel 9, released in January 2026, removed the old WORKSPACE system for declaring dependencies entirely, so projects that never migrated are now stuck on older versions.

Most teams can get most of the value more cheaply. Pin every dependency. Refer to container images by their digest, a hash of their contents, rather than a tag that can move. Mirror every registry you pull from. Generate a software bill of materials, a list of every component in what you ship. And build each artifact once, then promote that same artifact from environment to environment, so production runs exactly what was tested.

Reproducibility isn’t the whole story, though. The xz backdoor in 2024 was hidden in release tarballs and never appeared in git, and the malicious tarball built perfectly reproducibly. What catches that is provenance: a signed record linking an artifact to the exact source it came from. SLSA, the main framework for supply-chain security, is built around provenance and doesn’t require reproducibility at all.

What you depend on

Every tool you adopt is a bet on the people who maintain it, and the last few years have tested a lot of those bets. Vault, Terraform, and Nomad moved to a source-available license in 2023, and the community forked Vault as OpenBao. Ingress NGINX, for years the default way into a Kubernetes cluster, was retired in March 2026. Nix went through a governance crisis, and in August 2026 its core packaging team disbanded.

Each turned a stable dependency into a migration someone had to schedule. The practical response is to own your mirrors, pin by digest so nothing changes underneath you, prefer projects with a foundation and several maintainers, and keep the whole surface small enough that any single piece could be replaced.

Orchestration

The idea worth keeping from modern orchestration is the reconcile loop. You declare what should be running, and a controller continuously compares that to what is running and corrects the difference. Kubernetes is the dominant implementation, but the loop is the idea, and it is what makes a system self-healing rather than a pile of scripts that worked once.

On your own hardware, Kubernetes itself is rarely the hard part. The hard part is everything a cloud provider normally supplies underneath it: load balancers, persistent storage, operating system patches, and cluster upgrades. Immutable operating systems like Talos, which have no shell and are managed entirely through an API, remove a whole class of machines slowly drifting apart. At the other end of the debate, 37signals dropped Kubernetes entirely for Kamal, which deploys containers over SSH. Simple setups like that work well until several teams need to deploy on their own.

Compute for training

Training is a different workload from serving. A service wants many small, long-lived processes that can be restarted anywhere. A distributed training job wants many large machines at once, all or nothing, for hours or days, and a single failed node can stall the whole run.

Research clusters have long used Slurm, the batch scheduler from high-performance computing, which queues jobs and hands out whole allocations. Kubernetes was built for services and has had to grow the same ideas, gang scheduling and fair-share queues, through projects like Kueue. Ray sits above either one: a Python-native framework that lets a single program spread its work across a cluster, run on Kubernetes through KubeRay.⁠The appeal of Ray is that the same code runs on a laptop and on a cluster. KubeRay gives each job or team its own Ray cluster on shared Kubernetes hardware, which trades some efficiency for isolation. Whatever the scheduler, the hard problems are the same: keeping expensive accelerators busy, sharing them fairly between teams, and checkpointing often enough that a failure costs minutes instead of days.

How change is recorded

GitOps keeps the desired state of every environment in git and runs an agent inside the cluster that applies it. The git history becomes the change record: who changed what, when, and with whose review, which is exactly the evidence a regulated organization needs.⁠The two main GitOps tools are Argo CD and Flux. Flux survived the 2024 shutdown of Weaveworks, the company that created it, because it had already been donated to the CNCF and its maintainers found new homes.

Two rules keep that record honest. Promote the same artifact, by digest, from staging to production, and never rebuild in between. And remember that rolling back code doesn’t roll back a database. Schema changes have to go forward only, in expand-then-contract steps: add the new structure alongside the old, deploy code that works with both, move the data, and remove the old structure in a later release.

Platforms

A platform is the decision to solve the path from merge to production once, for everyone, instead of once per team.

I led the architecture and engineering of a centralized deployment platform, built from scratch on top of the existing clusters. It gives every team and every kind of service one path from a repository to a running, reachable service, in the spirit of Vercel rather than raw cluster configuration. It owns everything engineers used to stitch together by hand: building a commit into an artifact, scheduling it onto the cluster, wiring up its address and credentials, and making every deploy observable and reversible.

The hard part of any platform is deciding where the abstraction ends. Hide too little and every team still has to understand the cluster; hide too much and the first service that doesn’t fit has to leave the platform entirely. Good platforms make the common case need almost no configuration and give everything else an escape hatch one layer down, without leaving the platform. That matters more with AI in the loop. An agent can write a service in minutes, but it can only ship one safely if the path to production is a single, deterministic process that checks every change the same way.

What’s open

Stateful services and migrations. Keeping a mirrored, pinned world up to date without reintroducing the supply-chain risk the mirrors exist to prevent. Scheduling scarce accelerators fairly across training, evaluation, and serving. And verification that can keep pace when AI agents write and ship a growing share of the code. Working on this on a small team made the underlying point clear: an organization can only move as fast as its infrastructure lets it change, check, and undo what it runs.