We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
From the blog
OpenAI Rate Limits in Elixir: Retries, Backoff and Fallbacks in Phoenix
By Liam Killingback ·
OpenAI Rate Limits in Elixir: Retries, Backoff and Fallbacks in Phoenix
The first AI feature in a Phoenix app usually works beautifully in development. Then a customer uploads forty documents at once, a batch job fans out across every Oban queue, and the logs fill with 429 Too Many Requests. Some calls fail, some hang, and a few succeed twice because something retried a request that had already been billed.
OpenAI rate limits in Elixir are not hard to handle, but the defaults will not handle them for you. This guide builds a small, production-shaped client with Req: retries that know the difference between “slow down” and “you are out of money”, backoff that reads OpenAI’s own reset headers, a client-side limiter so you stop sending requests you know will fail, Oban jobs that snooze instead of burning attempts, a fallback chain across models and providers, a circuit breaker, and tests that prove all of it without touching the network.
Everything here is plain Elixir. No SDK is required.
What an OpenAI rate limit actually looks like
OpenAI limits each organisation and project on several axes at once:
- Requests per minute (RPM) and tokens per minute (TPM), per model. A handful of huge prompts can exhaust TPM while RPM is nowhere near its ceiling.
- Batch and daily limits on some tiers and endpoints.
- Spend limits and prepaid balance, which surface as the same HTTP status code but mean something completely different.
When you cross a limit you get an HTTP 429 with a JSON body like this:
{
"error": {
"message": "Rate limit reached for gpt-4.1-mini on tokens per min...",
"type": "tokens",
"code": "rate_limit_exceeded"
}
}
And when the account has run out of credit, you get a 429 as well:
{
"error": {
"message": "You exceeded your current quota, please check your plan and billing details.",
"type": "insufficient_quota",
"code": "insufficient_quota"
}
}
The first one fixes itself in seconds. The second one never fixes itself until a human adds a card. Retrying it is pure waste, and retrying it with exponential backoff across a few hundred jobs is how a billing problem becomes an outage that lasts all afternoon.
Every response also carries headers that tell you where you stand:
| Header | Example | Meaning |
|---|---|---|
x-ratelimit-limit-requests |
500 |
RPM ceiling for this model |
x-ratelimit-remaining-requests |
0 |
Requests left in the current window |
x-ratelimit-reset-requests |
1s |
Time until the request window resets |
x-ratelimit-remaining-tokens |
1830 |
Tokens left in the current window |
x-ratelimit-reset-tokens |
6m0s |
Time until the token window resets |
The reset values are Go-style durations (20ms, 1s, 6m0s), not seconds. We will parse them.
Beyond 429s, you will also see the usual transient failures of any HTTP API: 500, 502, 503 and 504 from upstream, connection resets, and timeouts on long completions. If you also call Anthropic, add its 529 “overloaded” status to the list.
Req does not retry your OpenAI calls by default
This catches almost everyone. Req ships with a retry step, and its default mode is :safe_transient, which retries only GET and HEAD requests. Every OpenAI call is a POST, so out of the box a 429 on /chat/completions is returned to you on the first attempt with no retry at all.
You can switch to retry: :transient, which retries any method on 408, 429, 500, 502, 503 and 504. That is a real improvement, and for a prototype it is enough. It still has three problems:
-
It retries
insufficient_quotaexactly like a normal rate limit. - It does not treat Anthropic’s 529 as transient.
-
Its backoff is 1s, 2s, 4s, 8s unless the server sends
Retry-After. OpenAI often tells you precisely when the window resets inx-ratelimit-reset-*, and waiting 1 second for a token window that resets in 40 seconds just earns you three more 429s.
So we write our own retry function. Req makes this pleasant: :retry accepts a two-arity function that receives the request and either the response or the exception, and returns true, false or {:delay, milliseconds}.
A retry policy that knows the difference
Here is the whole client. Read the comments, then we will go through the parts that are easy to get wrong.
defmodule MyApp.AI.OpenAI do
@moduledoc "A small OpenAI client with rate-limit aware retries."
require Logger
@base_url "https://api.openai.com/v1"
@default_model "gpt-4.1-mini"
@transient [408, 409, 500, 502, 503, 504]
# Waits longer than this are not worth blocking a request on.
# Return an error and let the caller (or Oban) come back later.
@max_inline_wait_ms 20_000
def chat(messages, opts \\ []) do
body = %{model: Keyword.get(opts, :model, @default_model), messages: messages}
case Req.post(client(), url: "/chat/completions", json: body) do
{:ok, %Req.Response{status: 200, body: body}} ->
{:ok, get_in(body, ["choices", Access.at(0), "message", "content"])}
{:ok, %Req.Response{status: 429} = resp} ->
{:error, {:rate_limited, error_code(resp), wait_ms(resp)}}
{:ok, %Req.Response{status: status, body: body}} ->
{:error, {:http, status, body}}
{:error, exception} ->
{:error, exception}
end
end
def client do
Req.new(
base_url: @base_url,
auth: {:bearer, Application.fetch_env!(:my_app, :openai_api_key)},
# The retry step runs before the body is decompressed or decoded.
# Asking for an uncompressed body lets retry?/2 read the error code.
compressed: false,
receive_timeout: 90_000,
retry: &retry?/2,
max_retries: 4
)
|> Req.merge(Application.get_env(:my_app, :openai_req_options, []))
end
@doc false
def retry?(_request, %Req.Response{status: 429} = resp) do
case error_code(resp) do
"insufficient_quota" ->
Logger.error("OpenAI quota exhausted: not retrying")
false
_rate_limited ->
case wait_ms(resp) do
nil -> true
ms when ms <= @max_inline_wait_ms -> {:delay, ms + :rand.uniform(250)}
_too_long -> false
end
end
end
def retry?(_request, %Req.Response{status: status}) when status in @transient, do: true
def retry?(_request, %Req.TransportError{reason: reason})
when reason in [:timeout, :econnrefused, :closed],
do: true
def retry?(_request, _response_or_exception), do: false
# How long OpenAI says to wait: Retry-After if present, otherwise the
# reset header for whichever window is exhausted.
defp wait_ms(resp) do
case Req.Response.get_header(resp, "retry-after") do
[seconds | _] ->
String.to_integer(seconds) * 1_000
[] ->
[{"requests", "x-ratelimit-reset-requests"}, {"tokens", "x-ratelimit-reset-tokens"}]
|> Enum.filter(fn {kind, _} -> exhausted?(resp, kind) end)
|> Enum.flat_map(fn {_, header} -> Req.Response.get_header(resp, header) end)
|> Enum.map(&parse_duration/1)
|> Enum.max(fn -> nil end)
end
end
defp exhausted?(resp, kind) do
case Req.Response.get_header(resp, "x-ratelimit-remaining-#{kind}") do
[remaining | _] -> String.to_integer(remaining) == 0
[] -> false
end
end
# "20ms", "1s", "6m0s", "1h2m3.5s" -> milliseconds
defp parse_duration(value) do
~r/(\d+(?:\.\d+)?)(ms|s|m|h)/
|> Regex.scan(value)
|> Enum.reduce(0, fn [_, n, unit], acc ->
{n, _} = Float.parse(n)
acc + n * Map.fetch!(%{"ms" => 1, "s" => 1_000, "m" => 60_000, "h" => 3_600_000}, unit)
end)
|> round()
end
defp error_code(%Req.Response{body: %{"error" => %{"code" => code}}}), do: code
defp error_code(%Req.Response{body: body}) when is_binary(body) do
case Jason.decode(body) do
{:ok, %{"error" => %{"code" => code}}} -> code
_ -> nil
end
end
defp error_code(_resp), do: nil
end
Why compressed: false
Req’s response pipeline runs the retry step before it decompresses and decodes the body. By default Req asks for a gzip response, so inside retry?/2 the body can be compressed bytes, and your careful check for insufficient_quota silently never matches. Turning compression off makes the body a plain JSON string at that point, which error_code/1 decodes. Chat completion responses are small enough that the bandwidth cost is irrelevant. After the pipeline finishes, the body is a decoded map again, which is why error_code/1 has a clause for both shapes.
Why a cap on the inline wait
A retry inside Req.post blocks the calling process. In a LiveView or a controller, that is a user staring at a spinner. Waiting two seconds for the request window is fine. Waiting six minutes for the token window is not: at that point the honest answer is an error the UI can show (“We’re busy, try again in a few minutes”) or a job that comes back later. So anything longer than @max_inline_wait_ms stops retrying and returns {:error, {:rate_limited, code, wait_ms}} with the wait attached, which the caller can use.
Why jitter
If fifty processes all hit the limit at the same moment and all wait exactly reset + 0ms, they all come back at the same moment and hit it again. A few hundred milliseconds of random spread is enough to break that up.
Timeouts are the uncomfortable case
Retrying a :timeout is not free. A long completion that timed out on your side may have finished on OpenAI’s side, which means you paid for it and will now pay again. For short chat calls this is an acceptable cost. For expensive calls (long documents, reasoning models), consider raising receive_timeout well above the default 15 seconds, as we did, and not retrying timeouts at all. If you bill customers for AI usage, meter the tokens the provider actually reports rather than the number of attempts; Aurora Meter exists for exactly that job.
Reading the headers before you hit the wall
Retries are the reactive half. The headers let you be proactive. A tiny Req response step can emit telemetry on every call, so you can chart how close you run to the ceiling and alert before customers see errors:
defmodule MyApp.AI.RateLimitTelemetry do
def attach(%Req.Request{} = request) do
Req.Request.append_response_steps(request, rate_limit_telemetry: &emit/1)
end
defp emit({request, response}) do
measurements =
for kind <- ["requests", "tokens"],
[value | _] <- [Req.Response.get_header(response, "x-ratelimit-remaining-#{kind}")],
into: %{},
do: {String.to_atom("remaining_" <> kind), String.to_integer(value)}
:telemetry.execute([:my_app, :openai, :rate_limit], measurements, %{
status: response.status
})
{request, response}
end
end
Pipe the client through it (client() |> MyApp.AI.RateLimitTelemetry.attach()), add a last_value metric in your Telemetry module, and LiveDashboard will draw the remaining budget per window. If remaining_tokens regularly dips under ten percent during business hours, you need either a higher tier or the next section.
Stop sending what you know will fail: a client-side limiter
Every 429 is a round trip you paid latency for. If your app can generate more concurrent requests than your tier allows (a bulk import, a popular feature at 9am), cap concurrency before the request leaves the node. ETS gives you an atomic counter for free:
defmodule MyApp.AI.Limiter do
@table __MODULE__
# Call once from Application.start/2, before the supervisor starts.
def init do
:ets.new(@table, [:named_table, :public, write_concurrency: true])
:ets.insert(@table, {:in_flight, 0})
:ok
end
def run(max_in_flight, fun) do
if :ets.update_counter(@table, :in_flight, 1) > max_in_flight do
:ets.update_counter(@table, :in_flight, -1)
{:error, :busy}
else
try do
fun.()
after
:ets.update_counter(@table, :in_flight, -1)
end
end
end
end
MyApp.AI.Limiter.run(20, fn ->
MyApp.AI.OpenAI.chat(messages)
end)
:ets.update_counter/3 is atomic, so two processes can never both take the twentieth slot. Two honest caveats. First, this is per node: on a three-node cluster you are allowing sixty in flight, so divide your budget by the node count or accept some 429s. Second, if the calling process is killed with Process.exit(pid, :kill) mid-call, the after block does not run and a slot leaks until restart. For a user-facing path that is rare enough to live with. For heavy batch work, the better limiter is a queue, which brings us to Oban.
Batch work belongs in Oban, and 429s should snooze
If you are summarising a thousand documents, do not loop over them in a request. Put each one in an Oban job on a dedicated queue, and let the queue’s concurrency be your rate limiter:
# config/config.exs
config :my_app, Oban,
repo: MyApp.Repo,
queues: [default: 10, ai: 5]
Then teach the worker what each error means:
defmodule MyApp.Workers.SummariseDocument do
use Oban.Worker, queue: :ai, max_attempts: 5
alias MyApp.AI.OpenAI
@impl Oban.Worker
def perform(%Oban.Job{args: %{"document_id" => id}}) do
doc = MyApp.Documents.get!(id)
messages = [%{role: "user", content: "Summarise this document:\n\n" <> doc.body}]
case OpenAI.chat(messages) do
{:ok, summary} ->
MyApp.Documents.save_summary!(doc, summary)
:ok
{:error, {:rate_limited, "insufficient_quota", _wait}} ->
# No amount of waiting fixes this. Stop, and alert a human.
{:cancel, :insufficient_quota}
{:error, {:rate_limited, _code, wait_ms}} ->
# Come back when the window has reset, without counting it as a failure.
{:snooze, div(wait_ms || 30_000, 1_000) + 1}
{:error, reason} ->
{:error, reason}
end
end
end
The three return values carry three different meanings. {:snooze, seconds} reschedules the job without treating it as a failure, so a busy afternoon does not eat into max_attempts. {:cancel, reason} stops the job permanently, which is right for a billing problem. {:error, reason} is a genuine failure and gets Oban’s normal backoff. If you have not used Oban much, our practical guide to Oban background jobs in Phoenix covers queues, uniqueness and cron in more depth.
One thing to know: queue concurrency in open source Oban is per node, like the ETS limiter. Oban Pro adds global and rate-limited queues if you need an exact cluster-wide ceiling.
Fallbacks: another model, or another provider
Sometimes the right response to a rate limit is not to wait but to go somewhere else. A cheaper model on the same account usually has its own RPM and TPM budget, and a second provider has an entirely separate one. A fallback chain is a reduce_while:
defmodule MyApp.AI do
@chain [
{MyApp.AI.OpenAI, model: "gpt-4.1-mini"},
{MyApp.AI.OpenAI, model: "gpt-4.1-nano"},
{MyApp.AI.Anthropic, model: "claude-haiku-4-5"}
]
def chat(messages) do
Enum.reduce_while(@chain, {:error, :no_providers}, fn {provider, opts}, _last ->
case provider.chat(messages, opts) do
{:ok, _text} = ok -> {:halt, ok}
{:error, reason} = error -> if fallback?(reason), do: {:cont, error}, else: {:halt, error}
end
end)
end
# Capacity problems move on to the next provider. Our own bugs do not.
defp fallback?({:rate_limited, _code, _wait}), do: true
defp fallback?({:http, status, _body}) when status >= 500, do: true
defp fallback?(:circuit_open), do: true
defp fallback?(%Req.TransportError{}), do: true
defp fallback?(_other), do: false
end
The key design decision is in fallback?/1. A 400 means your request is malformed (a bad schema, a missing field), and sending the same broken request to three providers only triples the latency before you see the error. Only capacity problems should fall through.
Be honest with yourself about the trade: models are not interchangeable. A fallback that returns a noticeably worse answer, or JSON in a slightly different shape, can be worse than a clear “try again”. Keep each provider module responsible for returning the same normalised result (plain text here, or a validated struct for structured output), and run your evals against every model in the chain. Our post on testing AI features in Phoenix with ExUnit shows how to build those.
A circuit breaker for when the provider is down
Retries and fallbacks assume the provider is basically healthy. During a real incident, every request spends its full retry budget failing before it falls back, and your users wait through all of it. A circuit breaker notices the pattern and skips the doomed provider for a while. The Erlang fuse library does this in a few lines:
# In Application.start/2: blow after 5 failures in 10 seconds, reset after 30 seconds
:fuse.install(:openai, {{:standard, 5, 10_000}, {:reset, 30_000}})
def chat(messages, opts \\ []) do
case :fuse.ask(:openai, :sync) do
:blown ->
{:error, :circuit_open}
:ok ->
result = do_chat(messages, opts)
if server_side_failure?(result), do: :fuse.melt(:openai)
result
end
end
defp server_side_failure?({:error, {:http, status, _}}) when status >= 500, do: true
defp server_side_failure?({:error, %Req.TransportError{}}), do: true
defp server_side_failure?(_result), do: false
Notice that 429s do not melt the fuse. A rate limit means the provider is healthy and you are busy, which is a job for the limiter and the backoff, not for the breaker. Combined with the fallback chain, an OpenAI outage now costs one fast :circuit_open check per request instead of four slow failures.
Streaming changes the rules
If you stream tokens into a LiveView, as in our guide to streaming OpenAI responses in Phoenix LiveView, a retry is only safe before the first token arrives. A 429 comes back as the response status before any body, so the retry policy above still applies to it. But if the connection drops halfway through an answer, blindly retrying would append a second, different answer to the half that is already on screen.
The pattern that works: track whether any chunk has been received. If none has, retry or fall back as normal. If some have, stop, keep what the user has, mark the message as interrupted, and render a “Regenerate” button that starts a fresh request and replaces the partial text. Users understand an interrupted answer. They do not understand two answers glued together.
Testing it without calling OpenAI
Req ships with Req.Test, which lets you stub the HTTP layer with a plug. Because our client merges :openai_req_options from config, the test setup is one line:
# config/test.exs
config :my_app, :openai_api_key, "test-key"
config :my_app, :openai_req_options,
plug: {Req.Test, MyApp.AI.OpenAI},
retry_delay: fn _attempt -> 0 end
Req.Test.expect/3 queues responses in order, which is exactly what a retry test needs:
defmodule MyApp.AI.OpenAITest do
use ExUnit.Case, async: true
alias MyApp.AI.OpenAI
setup {Req.Test, :verify_on_exit!}
test "retries a rate limit and then succeeds" do
Req.Test.expect(OpenAI, fn conn ->
conn
|> Plug.Conn.put_status(429)
|> Req.Test.json(%{error: %{code: "rate_limit_exceeded"}})
end)
Req.Test.expect(OpenAI, fn conn ->
Req.Test.json(conn, %{choices: [%{message: %{content: "Hello"}}]})
end)
assert {:ok, "Hello"} = OpenAI.chat([%{role: "user", content: "Hi"}])
end
test "does not retry an exhausted quota" do
Req.Test.expect(OpenAI, 1, fn conn ->
conn
|> Plug.Conn.put_status(429)
|> Req.Test.json(%{error: %{code: "insufficient_quota"}})
end)
assert {:error, {:rate_limited, "insufficient_quota", _}} =
OpenAI.chat([%{role: "user", content: "Hi"}])
end
test "retries a dropped connection" do
Req.Test.expect(OpenAI, &Req.Test.transport_error(&1, :closed))
Req.Test.expect(OpenAI, &Req.Test.json(&1, %{choices: [%{message: %{content: "Back"}}]}))
assert {:ok, "Back"} = OpenAI.chat([%{role: "user", content: "Hi"}])
end
end
verify_on_exit! fails the test if an expected response was never requested, so the second test also proves the client made exactly one call. That is the property you care about: a quota problem costs one request, not five.
A checklist for production
-
Use a custom
:retryfunction. Req’s default never retries a POST. -
Never retry
insufficient_quota. Alert on it instead. - Wait for the window the headers name, plus jitter, and cap inline waits.
- Decide deliberately whether timeouts are retried, because a timed-out call may still be billed.
-
Emit the
x-ratelimit-remaining-*headers as telemetry and alert at ten percent. - Cap concurrency before the request leaves the node: ETS for request paths, an Oban queue for batches.
-
In Oban,
{:snooze, seconds}for rate limits and{:cancel, reason}for billing problems. - Fall back only on capacity errors, and eval every model in the chain.
- Trip a circuit breaker on 5xx and transport errors, not on 429s.
- Retry streams only before the first token.
-
Prove all of it with
Req.Test.
Summary
Handling OpenAI rate limits in Elixir comes down to being specific about failure. A 429 for “tokens per minute” wants a short, header-guided wait. A 429 for “insufficient quota” wants a human. A 503 wants a retry, a string of them wants a circuit breaker, and a batch of a thousand documents wants a queue that snoozes rather than a loop that hammers. The BEAM gives you every piece you need (atomic ETS counters, supervised jobs, cheap processes) and Req gives you a retry hook that sees the full response.
If you would rather start from a codebase that already has an AI boundary to put this policy behind, the PHX AI Document Starter is a Phoenix app for uploading PDFs and chatting with them through a provider-agnostic AI layer (OpenAI in production, a deterministic fake in development and tests). Drop the retry, limiter and fallback modules above into that boundary and every AI call in the app inherits them. And if you are planning to build more than one AI product, the Builder Pass gives you every PhxTemplates starter, current and future, for a single lifetime price.