We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
From the blog
AI Document Summarization in Elixir: Summarize Long PDFs with Oban and LiveView
By Liam Killingback ·
AI Document Summarization in Elixir: Summarize Long PDFs with Oban and LiveView
Summarising a two-page memo with a language model is one API call. Summarising a 180-page contract, a year of board minutes or a technical manual is not, and the naive version fails in ways that are easy to miss in a demo: the text does not fit in the context window, the model quietly skips the middle of what does fit, the request takes long enough to hit a timeout, and a retry pays for the whole document again.
This post builds AI document summarization in Elixir that holds up on long inputs. We will extract text from a PDF while keeping page numbers, split it into token-sized chunks, summarise the chunks in parallel, combine the results (recursively if they are still too long), cache every intermediate step so a retry is cheap, and stream progress into a Phoenix LiveView. It is part of our series on building AI apps with Elixir and Phoenix, and it pairs well with the post on extracting structured data from PDFs: extraction pulls fields out, summarization compresses everything else.
Why one big prompt does not work
There are three separate problems, and they compound.
Length. A 200-page document is often 100,000 tokens or more. Even when a model accepts that much input, you pay for every token on every attempt, and your output budget is small by comparison.
Attention. Long-context models are measurably better at the start and end of their input than the middle. A single-shot summary of a long report tends to be a good summary of the introduction and the conclusion with a vague paragraph in between. For a contract, the middle is where the indemnity clause lives.
Operations. A single request that runs for minutes is fragile. If it fails at 90% you have nothing to show for it, and the user has been staring at a spinner the whole time.
The fix is old and well understood: map-reduce summarization. Summarise each section on its own (map), then summarise the summaries (reduce). It turns one fragile, expensive call into many small, cacheable, parallel ones, and that shape is exactly what the BEAM is good at.
Three strategies, and when each one fits
| Strategy | How it works | Good for | Weakness |
|---|---|---|---|
| Stuff | Whole document in one prompt | Short documents (under ~10 pages) | Breaks on length, loses the middle |
| Map-reduce | Summarise chunks in parallel, then combine | Long documents, reports, manuals | Can lose threads that span chunks |
| Refine | Walk the chunks in order, updating a running summary | Narratives where order matters | Sequential, so slow; errors accumulate |
We will build map-reduce and fall back to “stuff” automatically when a document is short enough. Refine is worth knowing about, but on a 40-chunk document it means 40 sequential calls, and the parallelism is half the reason to do this in Elixir.
Step 1: Extract text and keep the page numbers
A summary that cannot point back to its source is hard to trust. Keep the page number attached to every piece of text from the start, so the final summary can cite pages.
pdftotext (from Poppler) separates pages with a form feed character, which makes this simple:
defmodule MyApp.Docs.Extract do
@doc "Returns {:ok, [%{page: n, text: text}]} for a text-based PDF."
def pages(path) do
case System.cmd("pdftotext", ["-layout", "-enc", "UTF-8", path, "-"]) do
{text, 0} ->
pages =
text
|> String.split("\f")
|> Enum.with_index(1)
|> Enum.map(fn {page_text, n} -> %{page: n, text: String.trim(page_text)} end)
|> Enum.reject(&(&1.text == ""))
{:ok, pages}
{output, code} ->
{:error, {:pdftotext_failed, code, output}}
end
end
end
Indexing before rejecting empty pages matters: a blank page 7 should not shift page 8 to “page 7”. If a PDF returns no text at all, it is a scan, and you need OCR or a vision model first. The PDF extraction post covers that route.
Step 2: Chunk by tokens, on paragraph boundaries
Chunks should be large enough that each one carries a coherent idea and small enough that the model reads all of it carefully. Around 2,000 to 4,000 tokens per chunk works well for summarization. You do not need an exact tokenizer for this: roughly four characters per token for English prose is close enough to size chunks, as long as you leave headroom.
Split on blank lines so a chunk never cuts a paragraph in half, and track the first and last page each chunk covers:
defmodule MyApp.Docs.Chunker do
@chars_per_token 4
def chunk(pages, max_tokens \\ 3_000) do
max_chars = max_tokens * @chars_per_token
pages
|> Enum.flat_map(fn %{page: page, text: text} ->
text
|> String.split(~r/\n\s*\n/, trim: true)
|> Enum.map(&{page, &1})
end)
|> Enum.chunk_while(
nil,
fn {page, para}, acc ->
cond do
acc == nil -> {:cont, new_chunk(page, para)}
acc.size + String.length(para) > max_chars -> {:cont, finish(acc), new_chunk(page, para)}
true -> {:cont, add(acc, page, para)}
end
end,
fn
nil -> {:cont, nil}
acc -> {:cont, finish(acc), nil}
end
)
end
defp new_chunk(page, para),
do: %{first: page, last: page, parts: [para], size: String.length(para)}
defp add(acc, page, para),
do: %{acc | last: page, parts: [para | acc.parts], size: acc.size + String.length(para)}
defp finish(acc),
do: %{pages: {acc.first, acc.last}, text: acc.parts |> Enum.reverse() |> Enum.join("\n\n")}
end
One edge case to handle before production: a single “paragraph” longer than the limit (a table extracted as one block, for example) becomes one oversized chunk. Add a fallback that slices any part over max_chars with String.slice/3 before it enters the chunker.
Step 3: A small, swappable model client
Keep the model call behind a behaviour. It costs five lines and it is what makes this testable without an API key, which we will need later.
defmodule MyApp.AI do
@callback complete(system :: String.t(), user :: String.t(), opts :: keyword()) ::
{:ok, String.t()} | {:error, term()}
def complete(system, user, opts \\ []), do: impl().complete(system, user, opts)
defp impl, do: Application.get_env(:my_app, :ai_client, MyApp.AI.OpenAI)
end
defmodule MyApp.AI.OpenAI do
@behaviour MyApp.AI
@impl true
def complete(system, user, opts) do
body = %{
model: Keyword.get(opts, :model, config(:map_model)),
temperature: 0.2,
messages: [
%{role: "system", content: system},
%{role: "user", content: user}
]
}
case Req.post("https://api.openai.com/v1/chat/completions",
json: body,
auth: {:bearer, config(:api_key)},
receive_timeout: 120_000,
retry: :transient
) do
{:ok, %{status: 200, body: %{"choices" => [%{"message" => %{"content" => text}} | _]}}} ->
{:ok, text}
{:ok, %{status: status, body: body}} ->
{:error, {:http, status, body}}
{:error, reason} ->
{:error, reason}
end
end
defp config(key), do: Application.fetch_env!(:my_app, :openai) |> Keyword.fetch!(key)
end
retry: :transient tells Req to retry a POST on 408, 429 and 5xx responses, which it will not do by default. For a proper policy with backoff that reads the rate-limit headers, see OpenAI rate limits in Elixir.
Two models are configured on purpose: a small, cheap one for the map step (it runs dozens of times per document) and a stronger one for the final reduce (it runs once and its output is what the user reads).
# config/runtime.exs
config :my_app, :openai,
api_key: System.fetch_env!("OPENAI_API_KEY"),
map_model: "gpt-4o-mini",
reduce_model: "gpt-4o"
Use whichever models are current for you; the structure does not change.
Step 4: Map, in parallel, with a cache in front
The map prompt matters more than people expect. Tell the model what it is looking at, what to keep and what not to invent:
defmodule MyApp.Docs.Summariser do
alias MyApp.Docs.Cache
@prompt_version "v3"
@map_prompt """
You are summarising one section of a longer document. Write 4 to 8 bullet points
covering the facts, decisions, obligations, numbers and dates in this section.
Keep names, amounts, dates and defined terms exactly as written. Do not add
anything that is not in the text. If the section is boilerplate, say so in one line.
"""
def map_chunks(chunks, opts \\ []) do
progress = Keyword.get(opts, :progress, fn _done -> :ok end)
chunks
|> Task.async_stream(&summarise_chunk/1,
max_concurrency: 6,
timeout: 180_000,
on_timeout: :kill_task
)
|> Stream.with_index(1)
|> Enum.reduce_while({:ok, []}, fn
{{:ok, {:ok, section}}, done}, {:ok, acc} ->
progress.(done)
{:cont, {:ok, [section | acc]}}
{{:ok, {:error, reason}}, _done}, _acc ->
{:halt, {:error, reason}}
{{:exit, reason}, _done}, _acc ->
{:halt, {:error, {:chunk_exit, reason}}}
end)
|> case do
{:ok, sections} -> {:ok, Enum.reverse(sections)}
error -> error
end
end
defp summarise_chunk(%{pages: pages, text: text}) do
model = model(:map_model)
with {:ok, summary} <-
Cache.fetch([@prompt_version, model, text], fn ->
MyApp.AI.complete(@map_prompt, text, model: model)
end) do
{:ok, %{pages: pages, summary: summary}}
end
end
defp model(key), do: Application.fetch_env!(:my_app, :openai) |> Keyword.fetch!(key)
end
Task.async_stream/3 keeps results in input order, caps concurrency (six parallel calls is polite to most rate limits; raise it if your tier allows), and turns a hung request into an {:exit, :timeout} instead of a hung job. Halting on the first error is deliberate: the job will be retried, and the cache below means the chunks that already succeeded are not paid for twice.
The cache is what makes retries cheap
Key the cache on everything that changes the output: the prompt version, the model and the exact input text. Bump @prompt_version when you edit a prompt and old entries simply stop matching.
defmodule MyApp.Docs.Cache do
use Ecto.Schema
alias MyApp.Repo
@primary_key {:key, :string, autogenerate: false}
schema "summary_cache" do
field :text, :string
timestamps(updated_at: false)
end
def fetch(key_parts, fun) do
key = :crypto.hash(:sha256, Enum.join(key_parts, <<0>>)) |> Base.encode16(case: :lower)
case Repo.get(__MODULE__, key) do
%__MODULE__{text: text} ->
{:ok, text}
nil ->
with {:ok, text} <- fun.() do
Repo.insert(%__MODULE__{key: key, text: text}, on_conflict: :nothing, conflict_target: :key)
{:ok, text}
end
end
end
end
A side effect worth having: two users who upload the same PDF share every chunk summary. If your documents are tenant-private and that sharing matters to you, add the tenant id to the key parts.
Step 5: Reduce, recursively when it is still too long
Forty chunk summaries of a few hundred tokens each usually fit comfortably in one final call. Very long documents can produce more than that, so the reduce step collapses groups of summaries until the combined text fits, then writes the final summary:
# inside MyApp.Docs.Summariser
@reduce_budget_tokens 12_000
@collapse_prompt """
Combine these section summaries into one set of 6 to 10 bullet points.
Keep every page reference in the form [pp. X-Y]. Do not invent facts.
"""
@final_prompt """
Write a summary of the whole document for a busy reader.
Start with a three-sentence overview, then a "Key points" list, then a
"Numbers and dates" list. After each point, cite its source as [pp. X-Y]
using the page references given. Use only the information provided.
"""
def reduce([%{summary: _} | _] = sections) do
joined = Enum.map_join(sections, "\n\n", &format_section/1)
if div(String.length(joined), 4) > @reduce_budget_tokens do
with {:ok, collapsed} <- collapse(sections), do: reduce(collapsed)
else
MyApp.AI.complete(@final_prompt, joined, model: model(:reduce_model))
end
end
defp collapse(sections) do
sections
|> Enum.chunk_every(8)
|> Task.async_stream(&collapse_group/1, max_concurrency: 4, timeout: 180_000)
|> Enum.reduce_while({:ok, []}, fn
{:ok, {:ok, section}}, {:ok, acc} -> {:cont, {:ok, [section | acc]}}
{:ok, {:error, reason}}, _ -> {:halt, {:error, reason}}
{:exit, reason}, _ -> {:halt, {:error, {:collapse_exit, reason}}}
end)
|> case do
{:ok, acc} -> {:ok, Enum.reverse(acc)}
error -> error
end
end
defp collapse_group(group) do
{first, _} = hd(group).pages
{_, last} = List.last(group).pages
input = Enum.map_join(group, "\n\n", &format_section/1)
with {:ok, text} <-
Cache.fetch([@prompt_version, "collapse", input], fn ->
MyApp.AI.complete(@collapse_prompt, input, model: model(:map_model))
end) do
{:ok, %{pages: {first, last}, summary: text}}
end
end
defp format_section(%{pages: {a, b}, summary: s}), do: "[pp. #{a}-#{b}]\n#{s}"
Each round of collapsing divides the number of sections by eight, so even a very long document converges in two or three rounds. The page references survive every round because each collapsed section inherits the range of the group it came from.
For short documents, skip the map step entirely: if the whole text fits under the reduce budget, send it straight to the final prompt as a single section. That is the “stuff” strategy, chosen automatically.
Step 6: Run it in Oban, not in the LiveView process
A long document takes anywhere from twenty seconds to a few minutes. That work belongs in a background job that survives a closed tab, a deploy and a crash. If you have not set Oban up yet, the Oban guide covers it.
defmodule MyApp.Docs.SummariseWorker do
use Oban.Worker,
queue: :ai,
max_attempts: 5,
unique: [period: 600, keys: [:document_id]]
alias MyApp.Docs
alias MyApp.Docs.{Chunker, Extract, Summariser}
@impl Oban.Worker
def perform(%Oban.Job{args: %{"document_id" => id}}) do
doc = Docs.get_document!(id)
with {:ok, pages} <- Extract.pages(doc.path),
chunks = Chunker.chunk(pages),
total = length(chunks),
:ok <- broadcast(id, {:summary_progress, 0, total}),
{:ok, sections} <-
Summariser.map_chunks(chunks, progress: &broadcast(id, {:summary_progress, &1, total})),
{:ok, summary} <- Summariser.reduce(sections),
{:ok, doc} <- Docs.save_summary(doc, summary) do
broadcast(id, {:summary_ready, doc.summary})
end
end
@impl Oban.Worker
def timeout(_job), do: :timer.minutes(10)
defp broadcast(id, message),
do: Phoenix.PubSub.broadcast(MyApp.PubSub, "document:#{id}", message)
end
The unique option stops a double click from summarising the same document twice. Any {:error, reason} from the pipeline fails the attempt and Oban retries with backoff; on the retry every chunk that already succeeded comes straight out of the cache, so a failure at chunk 38 of 40 costs two calls, not forty.
Step 7: Stream progress into the LiveView
The LiveView only needs to enqueue the job and listen:
defmodule MyAppWeb.DocumentLive.Show do
use MyAppWeb, :live_view
alias MyApp.Docs
alias MyApp.Docs.SummariseWorker
def mount(%{"id" => id}, _session, socket) do
doc = Docs.get_document!(socket.assigns.current_user, id)
if connected?(socket), do: Phoenix.PubSub.subscribe(MyApp.PubSub, "document:#{doc.id}")
{:ok, assign(socket, doc: doc, progress: nil)}
end
def handle_event("summarise", _params, socket) do
{:ok, _job} = %{document_id: socket.assigns.doc.id} |> SummariseWorker.new() |> Oban.insert()
{:noreply, assign(socket, progress: {0, nil})}
end
def handle_info({:summary_progress, done, total}, socket),
do: {:noreply, assign(socket, progress: {done, total})}
def handle_info({:summary_ready, summary}, socket) do
{:noreply, socket |> assign(progress: nil) |> update(:doc, &%{&1 | summary: summary})}
end
end
<button :if={is_nil(@progress)} phx-click="summarise">Summarise this document</button>
<p :if={@progress && is_nil(elem(@progress, 1))}>Reading the document...</p>
<div :if={@progress && elem(@progress, 1)}>
<progress max={elem(@progress, 1)} value={elem(@progress, 0)}></progress>
<p>Summarised <%= elem(@progress, 0) %> of <%= elem(@progress, 1) %> sections</p>
</div>
<article :if={@doc.summary} class="prose">
<%= MyApp.Markdown.to_safe_html(@doc.summary) %>
</article>
Because the job broadcasts over PubSub, the progress bar works in every tab the user has open, and if they navigate away and come back, the finished summary is already saved on the document.
Step 8: Testing without calling the API
The behaviour from step 3 pays for itself here. A fake client that echoes a predictable summary lets you test chunking, ordering, caching and the recursive reduce without a network call:
defmodule MyApp.AI.Fake do
@behaviour MyApp.AI
@impl true
def complete(_system, user, _opts) do
{:ok, "- summary of #{byte_size(user)} bytes"}
end
end
# config/test.exs
config :my_app, ai_client: MyApp.AI.Fake
test "a retry reuses cached chunk summaries" do
chunks = for n <- 1..3, do: %{pages: {n, n}, text: "section #{n} " <> String.duplicate("x", 50)}
assert {:ok, first} = Summariser.map_chunks(chunks)
assert length(first) == 3
# Second run: same chunks, same answers, served from summary_cache
assert {:ok, ^first} = Summariser.map_chunks(chunks)
end
Task.async_stream runs each call in a child process, so to prove the second run made no calls, count calls with an Agent (or assert on the number of summary_cache rows) rather than relying on assert_received in the test process. The Ecto sandbox follows $callers, so the tasks share the test’s connection. Keep a separate, tagged eval suite with a handful of real documents and expected key points; the post on LLM evals in ExUnit shows how to run those without spending money on every mix test.
What it costs, roughly
You can estimate the bill before you run it. For a document of D tokens split into chunks of C tokens, with map summaries of about S tokens each:
-
Map input is about
Dtokens acrossD / Ccalls on the small model. -
Map output and reduce input are both about
(D / C) * Stokens. - The final output is a few hundred tokens on the larger model.
For a 100,000-token document with 3,000-token chunks and 300-token summaries, that is 34 map calls reading 100,000 tokens, and one reduce call reading about 10,000. The expensive model only ever sees a tenth of the document, which is the main reason this is cheaper than a single long-context call, as well as more reliable. Store the token counts from each response (the usage field) next to the summary and you will know your real numbers within a week.
Honest tradeoffs
- Threads that span chunks get thinner. If an argument builds over thirty pages, map-reduce sees thirty fragments. Larger chunks help; so does a refine pass over the map summaries for documents where narrative order matters.
- Summaries can still be wrong. Low temperature and “do not invent” instructions reduce hallucination; they do not eliminate it. The page citations are there so a reader can check, and you should tell users the summary is a starting point, not a substitute for reading a contract.
- Tables and scans need different handling. Text extraction mangles complex tables. If tables carry the meaning, extract them separately or use a vision model on those pages.
-
The cache stores document text derivatives. Treat
summary_cachewith the same retention and deletion rules as the documents themselves. When a user deletes a document, delete its cached summaries too, or key them so you can.
Skip the plumbing
Upload handling, per-user storage, a document viewer, a tested AI provider boundary with a fake for development, and chat over documents are already built in the PHX AI Document Starter. The map-reduce summariser above drops straight into its provider boundary, so you can spend your time on prompts and evaluation instead of upload forms and deletion lifecycles.
If you are likely to build more than one Phoenix product this year, the Builder Pass gives you every template, including the AI and SaaS starters, for one price.
Summary
AI document summarization in Elixir works best as map-reduce: extract text with page numbers, chunk on paragraph boundaries by approximate token count, summarise chunks in parallel with Task.async_stream, and combine the summaries, collapsing recursively when they are still too long. Cache every model call on its prompt version, model and input so retries cost almost nothing, run the pipeline in an Oban job, and push progress to LiveView over PubSub. Keep the model behind a behaviour so you can test the whole pipeline offline, and keep page citations all the way through so readers can verify what the model wrote.