What is Homa?
What does it bring to the table? What are its problems? And can I implement a pure Elixir client.
I saw an artcile in theregister and so I went through the four main sources I found, and also studied the current Homa Linux implementation. The 2021 paper and PlatformLab/Homa describe the ideas well, but neither is an accurate implementation specification for Homa in October 2026.
The quick answer
Homa is a message/RPC-oriented transport designed specifically for datacenter networks. It is trying to solve a different problem from TCP:
When thousands of machines are making very small RPCs to each other, how do you keep the 99th-percentile latency of the little messages extremely low without destroying throughput for large messages?
Its answer is roughly:
messages rather than streams + receiver-controlled transmission + shortest-message-first scheduling + network priority queues + no connections.
The results are impressive. In the 2021 Linux evaluation, Homa’s P99 latency for short messages was reported as roughly 19–72× lower than TCP and 7–83× lower than DCTCP under the evaluated high-load workloads. Even the median was several times better. USENIX
AS an Elixir developer, I think Homa is conceptually a remarkably good fit for BEAM. RPCs are independent messages, there is no stream ordering requirement between unrelated RPCs, and one socket can support huge numbers of independent conversations. That maps much more naturally to processes and messages than TCP connections do.
But there is an important catch that I’ll discuss in detail later, which rules out developing entirely in pure Elixir production Homa implementation.
What I might do is build a pure-Elixir Homa API with a very small native transport layer talking to the current Homa kernel module.
A truly pure-Elixir implementation of the wire protocol is possible, but as a fascinating project, and I think it would be worth building as a reference implementation. But it would probably throw away much of what makes Homa valuable.
First: forget TCP for a moment
The easiest way to understand Homa is to contrast the abstractions.
| TCP | Homa | |
|---|---|---|
| Fundamental abstraction | Byte stream | RPC/message |
| Connection | Stateful connection | Connectionless |
| Ordering | Strict byte ordering | Packets can arrive in any order |
| Multiplexing | Usually above TCP | Native independent RPCs |
| Congestion/scheduling | Primarily sender controlled | Receiver heavily controls transmission |
| Fairness goal | Traditionally flow fairness | Minimize message completion latency |
| Priority | Usually little semantic knowledge | Short messages favored |
| Receiver incast | Hard problem | Explicitly designed around it |
| Application sees | Stream | Complete request/response message |
Suppose machine A has these three messages ready:
RPC 1 2 KB
RPC 2 700 KB
RPC 3 8 KB
TCP doesn’t really understand that these are three independent units of work.
Homa does.
It tries to transmit approximately:
2 KB
8 KB
700 KB
rather than allowing packets from the 700 KB operation to obstruct the 2 KB operation.
This is SRPT: Shortest Remaining Processing Time. In Homa’s case, “processing” primarily means bytes remaining to transmit.
This is intentionally unfair to big messages over short periods. But statistically, SRPT is extremely good at minimizing completion time.
That matters enormously for RPC-based datacenter workloads.
Why the receiver is in charge
This is probably the most important Homa insight.
Imagine 100 servers suddenly sending a response to one machine.
Each sender knows:
"I have data to send."
But none of them knows:
"The receiver currently has 99 other machines trying to send to it."
The receiver does know that.
So Homa gives the receiver considerable control over which senders are permitted to transmit.
For larger messages, the basic conceptual interaction is:
Sender Receiver
"I have a 400 KB message"
───────────────►
decide when it should run
relative to everything else
GRANT bytes,
priority = N
◄───────────────
DATA
───────────────►
GRANT more
◄───────────────
DATA
───────────────►
This gives the receiver a global view of everything competing for its incoming link.
That is something TCP fundamentally doesn’t have.
The Homa work describes the combination of sender and receiver scheduling as approximating a distributed bipartite matching between machines that want to send and machines able to receive. The receiver can deliberately overcommit several senders so its link stays busy even if the highest-priority sender temporarily cannot transmit. Homa Transport
Short messages get an even faster path
Homa doesn’t want a tiny 500-byte RPC to require:
request permission
wait one RTT
send data
That would defeat the purpose.
Small messages can therefore be unscheduled: they go directly.
In the current 2026 implementation, Homa has evolved this distinction further.
A small/unscheduled message essentially does:
DATA ─────────────────────────►
A larger message requiring receiver scheduling now starts with:
START_MSG ────────────────────►
◄──────────────── GRANT
DATA ────────────────────►
This is actually a September 2026 protocol change. Large messages that need grants are now completely scheduled rather than sending an initial chunk of unscheduled data. START_MSG was added specifically to announce them. GitHub
What the priority queues are doing
Imagine two packets reach a switch:
700 KB RPC: packet 151
2 KB RPC: packet 1
If they simply sit in FIFO order, the tiny RPC can spend its life behind packets belonging to large transfers.
Homa exploits hardware priority queues in datacenter switches.
Conceptually:
PRIORITY 7 tiny RPCs
PRIORITY 6
PRIORITY 5
...
PRIORITY 1 large RPCs
The receiver determines priorities dynamically for scheduled traffic.
And the important result from the research is that Homa doesn’t need hundreds of priority levels. The ATC paper found surprisingly good results with only a few, with useful benefits even with two. USENIX
This is where Homa stops being “a better UDP protocol” and starts being a datacenter architecture.
You really want control of the fabric.
Reliability is interesting too
Homa is not unreliable like UDP.
It implements reliable RPC semantics, but does so around messages instead of an ordered stream.
Suppose packets representing:
bytes 0..1499
bytes 1500..2999
bytes 3000..4499
arrive as:
3000..4499
0..1499
That’s fine.
The receiver knows the message has a hole.
If the missing portion doesn’t arrive, it sends a RESEND.
There are also protocol messages such as:
DATA
GRANT
RESEND
BUSY
UNKNOWN
CUTOFFS
NEED_ACK
ACK
START_MSG
BUSY is particularly clever. A sender receiving a retransmission request may effectively say:
I haven’t died; I’m deliberately busy sending something more important.
That prevents the receiver from treating scheduling delay as failure.
And Homa provides at-most-once RPC semantics by retaining server-side state until it knows the client received the response. ACKs can be piggybacked or sent explicitly.
So this is considerably more sophisticated than:
UDP + retry
Homa References
| Source | What I would use it for | Warning |
|---|---|---|
| Homa Wiki | Best conceptual introduction | Some descriptions lag current code |
| ATC 2021 paper | Why Homa works and performance evidence | 2021 protocol/implementation snapshot |
PlatformLab/Homa | Architecture of a userspace implementation | Abandoned/stale |
go-homa | Example language binding to Linux Homa | Now significantly stale |
The old PlatformLab/Homa repository is especially interesting for Elixir. It explicitly separated the implementation into:
Transport
│
▼
Packet Driver
│
├── DPDK
└── potentially another driver
The transport implemented Homa and the packet driver provided unreliable packet I/O. It was intended to run entirely in userspace, bypassing the kernel. But that repository itself says development stopped and recommends the kernel implementation instead. GitHub
Architecturally beautiful.
go-homa is not a Go implementation of Homa.
It’s essentially:
Go API
│
syscalls / ioctl / sendmsg / recvmsg
│
Linux Homa kernel module
│
network
And its source is refreshingly small as a result.
The socket opens:
socket(AF_INET, SOCK_DGRAM, IPPROTO_HOMA)
and uses custom sendmsg, recvmsg, setsockopt, buffer management and ioctls.
But its latest commits are from February 2024.
That is now a serious problem.
For example, go-homa contains:
IPPROTO_HOMA = 0xFD # 253
because Homa had not yet received an official IP protocol number.
The current Homa kernel interface says:
IPPROTO_HOMA = 146
Homa received its IANA assignment in October 2024. GitHub
That’s just the obvious incompatibility.
The ABI changed more substantially in March 2025.
Old go-homa expects:
HOMA_RECVMSG_REQUEST
HOMA_RECVMSG_RESPONSE
HOMA_RECVMSG_NONBLOCKING
The current interface removed those request/response flags and introduced private RPC support using:
HOMA_SENDMSG_PRIVATE
The current homa_sendmsg_args also gained fields that go-homa doesn’t have. GitHub
Now the key question: can we make this pure Elixir?
There are two quite different meanings of “pure Elixir”.
Pure Elixir binding to the Linux Homa module
This is where I hit a fundamental problem.
OTP’s modern :socket can open sockets with an integer protocol, so this part is plausible:
:socket.open(:inet, :dgram, 146)
The Erlang socket API can also do native socket options and sendmsg/recvmsg. Erlang.org
At first glance this looks promising.
Unfortunately, Homa’s userspace ABI does two very un-BEAM-like things.
First, its send API effectively passes a pointer to a native homa_sendmsg_args structure through msg_control while setting msg_controllen to zero. go-homa even comments explicitly on this unusual mechanism. The normal OTP sendmsg interface deals in ordinary ancillary/control messages; it does not expose this arbitrary pointer trick. Homa sendmsg API
Second, and worse, receiving requires registering a native userspace buffer pool with the kernel:
struct homa_rcvbuf_args {
uint64_t start;
uint64_t length;
};
start is literally an address in your process.
The kernel then puts received message data into pages within that region and returns offsets to those pages. GitHub
Pure Elixir cannot safely say:
Here is the stable machine address of this BEAM binary.
Please let the kernel DMA/write into it.
BEAM intentionally abstracts that sort of memory ownership away.
So I don’t think a robust client for the current kernel ABI can be literally 100% Elixir using public OTP APIs.
Could we instead implement the actual Homa protocol in Elixir?
Yes.
This is much more interesting.
Using a raw socket and I could implement something like:
Homa.Packet
Homa.RPC
Homa.Sender
Homa.Receiver
Homa.GrantScheduler
Homa.Reassembly
Homa.Retransmit
Homa.Ack
Homa.Pacer
Homa.Peer
And the BEAM process model would make parts of this wonderfully elegant.
You can almost see the supervision tree:
Homa.Transport
│
├── Socket
│
├── Sender
│
├── Receiver
│
├── GrantScheduler
│
├── RetransmitTimer
│
└── PeerSupervisor
├── Peer
├── Peer
└── Peer
An RPC can even correspond naturally to a process or lightweight state machine.
From a software-design standpoint, this could be beautiful.
But Homa isn’t primarily interesting because the state machine is elegant.
It is interesting because it delivers microsecond-scale tail latency at enormous packet rates under saturation.
And that’s where pure BEAM becomes questionable.
The ATC implementation already found that software packet processing was a major limitation: neither highly optimized Homa nor TCP could reach anything like full 25-Gbps bandwidth on very short-message workloads without consuming large amounts of CPU. Their 100-Gbps experiments required on the order of 18 cores in some configurations. USENIX
Now insert:
raw socket
→ BEAM scheduler
→ Erlang message
→ binary matching
→ process
→ state transition
→ Erlang message
→ packet encoding
→ syscall
into the packet hot path.
It would work.
It probably would not be Homa anymore in the performance sense.
The possible architecture I would build
I think there is a very nice middle ground:
Elixir
┌──────────────────────────────────────┐
│ Homa │
│ │
│ request/3 │
│ request_async/3 │
│ await/2 │
│ recv/2 │
│ reply/3 │
│ cancel/2 │
│ │
│ RPC/process ownership │
│ timeouts │
│ telemetry │
│ supervision │
└──────────────────┬───────────────────┘
│
tiny native boundary
│
┌──────────────────▼───────────────────┐
│ homa_native │
│ │
│ socket() / bind() │
│ register receive region │
│ sendmsg() │
│ recvmsg() │
│ reply │
│ abort │
│ buffer page ownership │
└──────────────────┬───────────────────┘
│
┌──────────────────▼───────────────────┐
│ Homa Linux kernel module │
│ │
│ grants / SRPT / pacing │
│ retransmission / ACKs │
│ packet priorities │
│ NIC interaction │
│ RSS/GRO/GSO/etc. │
└──────────────────┬───────────────────┘
│
NIC
That gives what I would call an Elixir-native Homa client, even though perhaps 500–1,000 lines at the very bottom are going to be C/Rust/Zig.
Everything semantically interesting to the Elixir programmer remains Elixir.
The native layer should be boring.
No RPC state machine.
No supervision.
No scheduling policy.
No clever threading.
Just:
BEAM value ⇄ Homa syscall ABI
I would initially copy received data into a BEAM binary, release the Homa bpages immediately, and measure it. Only if the copy appears significantly in the profile would I build resource-backed zero-copy binaries whose destructor returns Homa pages.
That keeps memory ownership sane. But I’ll revisit issue a bit later.
Homa fits BEAM surprisingly well above that line
Imagine:
{:ok, socket} = Homa.open(port: 4000)
{:ok, ref} =
Homa.request(
socket,
{"10.0.0.42", 9000},
<<"lookup", key::binary>>
)
receive do
{:homa, ^ref, response} ->
...
end
There needn’t be:
one connection
one socket
one process
for every server.
You might have:
one Homa socket
10,000 outstanding RPCs
10,000 Elixir callers
and simply use Homa’s 64-bit RPC identity/completion cookie to demultiplex completions.
That starts to feel a lot like BEAM itself:
send an independent message
don't care about unrelated ordering
receive completion later
rather than forcing BEAM’s messaging model through a TCP byte stream.
This idea I find a little compelling.
The strong points
The biggest attraction is tail latency. Homa was designed from scratch around completion time rather than around sustaining a reliable ordered byte stream, and its measured gains under congestion were dramatic. USENIX
The second is incast control. Making the receiver schedule large incoming messages is fundamentally sensible in a datacenter, because the receiver knows its own incoming contention.
Third, Homa removes per-connection state. A service talking to 20,000 other machines doesn’t require 20,000 transport connections. The Homa overview explicitly identifies connectionlessness and message-based communication as core design choices. Homa Transport
Fourth, unrelated RPCs don’t suffer TCP-style stream head-of-line blocking. Lost packet N from one large RPC doesn’t inherently stop unrelated RPC N+1 from completing.
Fifth, Homa is a particularly attractive match for request/response architectures. You aren’t pretending an RPC is a stream when it isn’t.
And finally, Homa is not dead research. The kernel module is actively changing right now; there were commits on October 6, 2026, and recent work includes RHEL support, Linux 7 compatibility work, a Homa qdisc, API evolution and upstreaming work. GitHub
The weak points
The price for all this is substantial.
Homa is not an Internet transport in the way TCP or QUIC is. The architecture assumes a controlled datacenter environment where you understand NICs, switches, queues and host software.
Its priority-queue dependency can collide with existing network QoS policy. The Homa project itself recognizes this as an operational concern. Homa Transport
And I’m cautious about the Wiki’s very strong claim that Homa eliminates core congestion. The ATC paper is more restrained: it says Homa addresses edge congestion but doesn’t explicitly solve congestion in the network core, and discusses techniques such as packet spraying as complementary work. USENIX
Receiver overcommitment also consumes buffering. This remains an area of research; Stanford’s newer SIRD work specifically attacks Homa’s buffer utilization problem while trying to retain nearly the same latency/throughput characteristics. Homa Transport
Deployment is currently specialized. The current project lists known support for Mellanox ConnectX-4/5/6 and Intel E810-class NICs rather than presenting itself as universally validated hardware support. GitHub
The API and even wire protocol are still evolving. In just the period relevant to go-homa, we’ve had:
2024 IANA protocol number / API evolution
2025 private RPC API and userspace ABI changes
2026 entirely scheduled large messages + START_MSG
That is significant churn for something I’d want at the bottom of a production stack. GitHub
And there is still a 1,000,000-byte maximum Homa message in the current userspace API. GitHub
What I might build
I would make this a staged project where the pure-Elixir protocol implementation and the useful production client reinforce each other rather than forcing us to choose immediately:
-
homa_wirein pure Elixir. Encode/decode every current Homa wire structure from the currenthoma_wire.h; golden-vector tests against C-generated packets; no network I/O yet. This would teach me the protocol and gives us invaluable diagnostics/Wireshark-style capabilities. -
A pure-Elixir reference transport. Raw socket, START_MSG/DATA/GRANT, basic RPCs, retransmission, ACK and reassembly. Explicitly call it a reference implementation, not the performance transport. Run it against the Linux implementation for interoperability.
-
homa_native. Implement only the current Linux UAPI fromhoma.h: create/bind socket, register buffer region, send/receive/reply/abort. Keep this boundary microscopic. -
homa_ex. Build the pleasant OTP-facing interface: synchronous request, async request, mailbox completion, cancellation, servers and Telemetry. Make the backend behaviour-pluggable soHoma.Transport.ReferenceandHoma.Transport.Linuxexpose the same interface. -
Benchmark the boundary. Compare Linux
cp_node, the Elixir/native client, TCP via:gen_tcp, and the pure-Elixir reference client. Measure P50/P95/P99/P99.9, message rate, CPU, scheduler utilization and response size distributions. Only then investigate zero-copy. -
Keep protocol-version compatibility explicit. I would generate/check the native ABI against current
homa.h, and treat current HomaModule as the compatibility authority.
Well, yes I can dream but if I get a bored week of so full of meetings I might knockout a readable Elixir specification of Homa plus a genuinely useful high-performance Elixir client.
My assessment
If the goal is:
“Can we port
go-homato Elixir?”
Yes, but go-homa is now sufficiently outdated that I would not port it. I would reimplement the small client against the 2026 HomaModule UAPI.
If the goal is:
“Can we implement Homa itself in Elixir?”
Also yes, and it would be an excellent technical project. But I would treat that version as a reference/research implementation, not claim comparable Homa performance.
If the goal is:
“Can Elixir become a first-class application environment for Homa?”
That is the project I think is genuinely compelling.
Homa’s semantics and BEAM’s semantics fit each other unusually well. Let Linux do the nanosecond/microsecond packet scheduling, NIC interaction, retransmission and buffer mechanics; let Elixir do what it excels at: millions of concurrent independent conversations, supervision, routing, timeouts and application-level messaging.
And I would specifically design the library so that this:
Elixir process
│
│ request
▼
Homa RPC
│
│ response
▼
Elixir message
is the central abstraction.
That could be a substantially nicer RPC substrate for BEAM services than pretending everything is an ordered TCP stream.
If you want, the logical next step is for me to turn this into a concrete architecture for homa_ex — modules, supervision tree, public API, NIF boundary, current 2026 ABI structures, and a staged implementation roadmap.
References
- John Ousterhout, A Linux Kernel Implementation of the Homa Transport Protocol, USENIX ATC 2021.
- Homa Transport wiki.
- PlatformLab/HomaModule — current Linux kernel module,
homa.h, and man pages (the 2026 specification). - PlatformLab/Homa — older userspace implementation; development stopped.
- dpeckett/go-homa — Go binding to an earlier kernel ABI; useful as a design example, not a current spec.
- homa.h — userspace ABI, including the IANA IP protocol assignment (October 2024).
- Erlang
:socket— native socket options andsendmsg/recvmsg. - Konstantinos Prasopoulos, Ryan Kosta, Edouard Bugnion, and Marios Kogias, SIRD: A Sender-Informed, Receiver-Driven Datacenter Transport Protocol, USENIX NSDI 2025.