Supervisors By Hand
2026-10-06
I just got home from Goatmire, and I saw an interesting talk by Frank Hunleth: “Nerves by hand”. The talk was about demystifying the Nerves toolchain and how you get an Elixir application compiled with a Linux kernel and run that on bare metal. I kind of knew how most of it worked at a very high level, but Frank made it clear that there is no black magic happening. This way of learning things is something that my brain is wired for. So, with that in mind, I wanted to give a shot at doing OTP supervisors by hand. They might seem like black magic to some people, but they are absolutely not. In this article I intend to show this.
The source code for this article can be found here.
In the article, all code blocks are loadable into the REPL on the right (or bottom on mobile). You should have access to nearly an entire BEAM running in your browser. If at any point you want to start from a clean slate, press the reset button next to the input and the BEAM will restart into a fresh state.
The Four Primitives
If you know the primitives, skip ahead to the supervisor section.
There are four primitives you need to know to build your own
supervisor: send/receive,
spawn,
monitor,
and link.
I will briefly explain them here, and then with just those four things
we will build our own supervisor.
Spawn
On the BEAM, a unit of concurrency is a “process.” They are
independent computations that run from start to finish, and have their
own memory. Nobody can fiddle in it besides the process itself. A
process can be created using the spawn
function.
spawn(fn -> IO.puts "I am #{inspect(self())}" end)The code above starts a process, makes it print “I am pid”, and then the process terminates. That is it. You can put as complex of a function in there as you like. The “pid” of the process is its unique address. There will never be another pid like it in the whole runtime of the BEAM. So no other process can ever have the same pid.
Note: There can be duplicate PIDs when the counter wraps around, but there will never be two processes alive with the same pid. See here for more information.
Send/Receive
Since each process has its own memory, and only it has access to it,
you need a way to let processes synchronize (i.e., communicate with each
other). The primitive for this is send.
Sending a message puts an arbitrary value in the process’ mailbox. Think
of your actual mailbox outside. People put stuff in, and you take it out
when you want to.
Taking out a message is done with the receive
primitive. The receive expression will look at the
processes’ mailbox and see if there is a message. If there is, the
clause in the expression is executed, and the message is dropped from
the mailbox.
In the code below a process is spawned and it will wait for 1 message. If the process receives the message it will print it, and terminate. So if you were to send 2 letters, only one will be read. The other one will disappear.
my_mailbox =
spawn(fn ->
receive do
{:letter, _subject, content, from} ->
IO.puts("I got a letter from #{from}, saying #{content}")
end
end)# Send a letter to my_mailbox
send(my_mailbox, {:letter, "Invoice", "Gimme 100 bucks", "Scammer"})Monitor
If you are sending letters, you might want to keep an eye on the
mailbox you send it to. If that mailbox is gone for some reason, you
should stop sending letters. This is where the monitor
primitive comes in. You can monitor any pid and get notified as soon as
it terminates. The notification, of course, is a message sent to the
process that called monitor. The result of calling
monitor on a process is a reference.
These are random unique values, much like process ids. They can be
generated cheaply, and they are unique in your virtual machine. They are
useful for figuring out which monitor sent a message about a process
dying, as we’ll see later.
my_mailbox =
spawn(fn ->
receive do
{:letter, _subject, content, from} ->
IO.puts("I got a letter from #{from}, saying #{content}")
end
end)
# Get notified about a termination.
ref = Process.monitor(my_mailbox)# Send a letter to my_mailbox
send(my_mailbox, {:letter, "Invoice", "Gimme 100 bucks", "Scammer"})
# Wait for the termination message.
receive do
{:DOWN, _ref, :process, pid, reason} ->
IO.puts("My mailbox died, with reason #{inspect(reason)}")
endLink
A more powerful variant of monitoring is linking. Let’s say a process
is sending letters to a specific mailbox, but as soon as it’s gone, the
sender has no purpose anymore. We should therefore tell the sender to
terminate when this happens. The mechanism to do this is called linking,
done with the link
primitive. When two processes are linked, and one of them dies, the
linked process will die as well. This is a transitive mechanism, meaning
that if you link A to B, B to C, and C to D, all these processes will
terminate as soon as one of them dies.
In the code below you can start a process pid, and a
second process other_pid will link to it. What we expect to
happen is that when pid terminates after 1 second,
other_pid should terminate as well. This can be checked
using the Process.alive? function in Elixir. It will return
true if the process is still alive, or false
when it has terminated.
Note: Processes exit with a reason. If this reason is
:normal, the linked processes do not terminate. This
makes sense, since you want to have linking in place for failure
handling, not orchestration.
pid = spawn(fn -> Process.sleep(1000); exit(:oopsie) end)
other_pid = spawn(fn -> Process.link(pid) ; Process.sleep(5000) end)Process.alive?(other_pid)In the terminal snippet above a process is created that will exit
after 3 seconds with the reason :oopsie. The linked process
is the process of the repl, and that will terminate as soon as
pid terminates.
Building a Supervisor
We can create processes, we can communicate with them, and we can
keep an eye on them to see if they’re still alive. Those are the only
ingredients you need for a supervisor. Because I’m writing this in the
Varberg train station, we’ll use the name Goatherder for
supervision, and all our processes in the tree are going to be
Goats.
Goat’s Behavior
In a GenServer
you have a lot of callbacks: handle_info,
handle_call, handle_cast and so on. We’ll keep
it simple and only have one: message. The function should
accept the message and the current state of the process, and then return
{:ok, state}. A simple Goat that refuses to do anything
could look like this.
defmodule Goatherder.HelloGoat do
def message(message, state) do
IO.puts("Nobody tells the goat to '#{inspect(message)}'")
{:ok, state}
end
endThis is just the behavior of our goat; it defines what our
HelloGoat does when it receives a message. It does not
actually create a process for it. But, we can already give it a test
spin.
Goatherder.HelloGoat.message("faint!", nil)Goat
For each goat we want to create in our herd, we need a process that
will represent it. This will be the Goat module. In the
Goat module, we have 2 functions. A
create_goat function, and a goat_loop
function. When we call create_goat with a module, a process
is spawned that will register itself under the name of the module. Then,
the goat_loop function is called to wait for the first
incoming message for our goat.
The goat_loop function keeps track of the goat’s state.
Each goat’s initial state is nil. In a real
GenServer, this can be changed in the init
callback.
defmodule Goatherder.Goat do
require Logger
@doc """
Given a goat module, creates a process that executes that goat's logic. The
name of the goat process is always the name of the module (i.e., we can only
have one of each goat running).
"""
def create_goat(goat) do
goat_pid =
spawn(fn ->
Process.register(self(), goat)
goat_loop(goat, nil)
end)
{:ok, goat_pid}
end
# listen for messages, and relay them to the goat implementation.
# Terminate on all values except `{:ok, term()}`.
defp goat_loop(goat, state) do
receive do
message ->
case goat.message(message, state) do
{:ok, state} ->
goat_loop(goat, state)
{:error, error} ->
Logger.error("Goat error: #{inspect(error)}")
end
end
end
endWe can create a Goat as follows.
{:ok, goat} = Goatherder.Goat.create_goat(Goatherder.HelloGoat)
send(goat, "faint!")The goat_loop function makes use of the
receive primitive. For each message that arrives in the
goat’s mailbox, the module’s message/2 function is called.
The message and the state are passed in, and we expect the module to
either return {:ok, state} or
{:error, reason}. Anything else will cause the goat process
to crash. If the message was well received, the loop waits for the next
incoming message.
Goatherder
We have discussed the mechanics to create a goat process, but we still have no supervisor for them. The main task of a supervisor is to monitor the goats, and recreate them as soon as they fall over.
What we want to end up with is a way to create a list of atoms that
are Goat implementation, such as our
HelloGoat. Then, ask our Goatherder to start them all up,
and if they crash, restart them. The mechanism behind this boils down to
the following. When we create a Goatherder with a list of
Goat modules, the following happens:
- Create a process for the
Goatherder. Inside that process, the following happens:- Spawn a
Goatprocess with the given behavior. - Monitor that goat process.
- Spawn a
The Goatherder process will then be notified as soon as
one of its goats crashes by means of a {:DOWN...} message
we saw before. If such a message arrives, the Goatherder
process will do the following.
- Figure out which
Goatcrashed. - Spawn a new
Goatwith the original module.
There are some bookkeeping details we need to do, however. Let’s look
at the creation of a herd, first. The create function will
call spawn_herd on the goats it has been given, and then
monitor the entire herd in an infinite loop.
A goat is spawned by creating a Goat process for the
module, and subsequently calling monitor on it, making sure
the herder is notified if the goat faints.
The return value from spawn_goat carries with it all the
information we need about our goats. The process id of the goat we just
spawned, the module that it is running, and a reference.
The final step is now monitoring the goats, and acting on a crashed goat. If a goat crashes, we just create it again.
The monitor function keeps calling itself, and waits for
it to receive a message about a fainted goat. Remember that
monitor returns a reference, so we can use that reference
here to determine which goat it was that fainted.
The reference is looked up in the herd state, and we
fetch the module that this goat was running. Using that, the herder
simply spawns a new copy of that goat and stores it in its state,
waiting for the next goat to faint.
defmodule Goatherder do
alias Goatherder.Goat
require Logger
def create(goats) do
herd_pid = spawn(fn ->
herd = spawn_herd(goats)
monitor(herd)
end)
{:ok, herd_pid}
end
defp spawn_herd(goats) do
for goat <- goats, into: %{} do
spawn_goat(goat)
end
end
defp spawn_goat(goat) do
Logger.info "Creating goat #{inspect goat}"
{:ok, goat_pid} = Goat.create_goat(goat)
goat_ref = Process.monitor(goat_pid)
{goat_ref, {goat_pid, goat}}
end
defp monitor(herd) do
herd =
receive do
{:DOWN, ref, :process, pid, reason} when is_map_key(herd, ref) ->
Logger.warning("Goat #{inspect(pid)} exited: #{inspect(reason)}")
{_goat_pid, goat} = Map.get(herd, ref)
{goat_ref, {goat_pid, goat}} = spawn_goat(goat)
herd
|> Map.delete(ref)
|> Map.put(goat_ref, {goat_pid, goat})
_ignored ->
herd
end
monitor(herd)
end
endCreating a Herd
At this point we have all our logic in place, and we can create our
herd. Make sure all the modules (Goatherder,
Goatherder.Goat, and Goatherder.HelloGoat have
been loaded into the repl.
We first create our list of goats, which now only includes the
HelloGoat. Then, we create a Goatherder by
calling create and passing our list of goats. This will
create all the processes and monitor them.
defmodule Goatherder.Example do
def init do
goats = [Goatherder.HelloGoat]
Goatherder.create(goats)
end
endTo actually run the herd, we can use the following code. First, let’s create our herd.
Goatherder.Example.initOnce the herd is created, we can verify that our HelloGoat process is alive.
goat_pid = Process.whereis(Goatherder.HelloGoat)
Process.alive?(goat_pid)To send a message to it, we use the send primitive.
send(goat_pid, :faint)And now, we can also kill our Goat and watch it spawn
back automatically.
Process.exit(goat_pid, :faint)To check if it’s responding, we’ll ask it to faint once more.
goat_pid = Process.whereis(Goatherder.HelloGoat)
send(goat_pid, :faint)That’s about it for supervisors. Of course, these are not functionally equivalent to the real supervisors. However, I hope they demystified the supervisors a bit more, and you can use them with more confidence.
Missing Features
This is by no means an equivalent implementation of the real
Supervisor, but it gives you an idea of what it takes to
build something. If you want to play around with this, a few things that
are still missing, but fundamental to supervisors are the following.
- Rate limiting the amount of restarts
- Different supervision strategies
- Cast/call messages
- An actual behavior
- …
You could fork the repo and try adding one of these features.