24/7 livestream encoders: failover, fencing and scaling from zero
How a 24/7 YouTube livestream platform keeps a stream alive when its encoder dies: epoch fencing, careful failover, a KEDA plan for 0 to 50 pods, 10GB uploads.

On this page
The livestream side of Wavesgroup does one thing that sounds trivial: it takes a video a customer uploaded and keeps it playing on a YouTube live stream, around the clock. Nearly every live on the platform is a 24/7 one. Each encoder runs one ffmpeg process that pushes one file, looped, to one RTMP key.
The trivial part ends the first time an encoder dies. If nobody picks the stream up, YouTube ends the broadcast and the customer's live is gone. If two encoders pick it up, both push to the same key and the stream breaks in a different way. Most of the work on this system has been about that gap: noticing a dead encoder quickly, without mistaking a slow database for a dead fleet, and making sure exactly one encoder is ever pushing.
This post covers the ownership and failover design as it runs today, the Kubernetes and KEDA setup it's being moved onto, and the resumable upload path that feeds it files up to 10 GB.
One pod, one ffmpeg, one live
The encoder is deliberately simple. One process owns one live:
ffmpeg -re -stream_loop -1 -i media.mp4 \
-c:v copy -c:a copy -f flv "rtmp://<ingest>/<stream-key>"Copy mode means no transcode, so a cheap node can carry the stream. That only works if the file is already shaped for live: the render service forces a fixed keyframe interval and AAC 48 kHz stereo on its output, because in copy mode whatever is in the file goes on air as it is. When a stream resumes after a handoff, the seek goes on the input side (-ss before -i). An output-side seek in copy mode produced zero frames. Between files, or while a replacement starts, the encoder pushes a black standby stream so the RTMP key stays alive.
Jobs reach encoders by push, not by a shared queue. The backend runs a reconcile loop every 15 seconds, picks an encoder for each live that needs one, and writes a dispatch command. Encoders poll for their commands every 3 seconds and then claim the job.
Ownership: owner, epoch, heartbeat
The first version used a classic lease: a row with an expiry, renewed every couple of seconds under a Postgres advisory lock, 6 seconds to live. It worked while encoders talked to the database directly. When they moved behind the backend's API, ownership became three fields on the job row instead:
owner_node: which encoder holds the live.owner_epoch: a fencing token that only ever goes up.last_heartbeat_at: refreshed by the owner every 20 seconds.
Claiming happens in a serializable transaction with FOR UPDATE SKIP LOCKED, and taking over from another owner is hard on purpose:
-- inside a SERIALIZABLE transaction
SELECT owner_node, owner_epoch, last_heartbeat_at
FROM encoder_jobs WHERE id = $1 FOR UPDATE SKIP LOCKED;
-- foreign owner? refuse (409) unless its heartbeat is > 180s old
-- AND the claimant is the node the backend designated
UPDATE encoder_jobs
SET owner_node = $me, owner_epoch = owner_epoch + 1, last_heartbeat_at = now()
WHERE id = $1;That rule came from an incident on 2026-07-29: an owner that restarted kept getting 409 forever, because nothing let anyone take a live over from a node that no longer existed. The two conditions together let a real takeover through, but not a race between two healthy nodes.

Fencing on the encoder
The epoch only helps if the encoder checks it. Right before spawning ffmpeg, the encoder sends an empty state patch carrying its epoch. A 409 means someone newer owns the live, so it releases with reason stale_epoch and never touches RTMP. It repeats that probe right after spawn and periodically during playback.
The kill path matters as much. On 2026-07-06 a split-brain killed a live: an ffmpeg process blocked on a dead RTMP socket ignored SIGTERM and kept pushing for about 16 minutes while another encoder owned the stream. Now SIGTERM escalates to SIGKILL after 5 seconds.
Failing over without false alarms
Detecting a dead encoder is easy. Not detecting dead encoders that aren't dead is the hard part, and most of the guards in the sweep exist because of a specific bad day:
- A heartbeat counts as stale after 60 seconds, and it must stay stale for two consecutive sweeps before failover, about 75 seconds in total. A "zombie" whose heartbeat moves but whose playback doesn't needs three sweeps.
- On 2026-06-30 a database outage made every heartbeat look stale at once, and the sweep failed over healthy lives. Now, if the ghost query itself takes more than 2 seconds, the sweep gives the database a 120-second grace period.
- If at least three lives and at least half of them go stale together, the sweep treats it as a backend or database problem, not a fleet of dead encoders, and does nothing.
- After a backend restart there is a 120-second startup grace, and a live gets at most three failover attempts.
The failover itself has an order that matters:
- Refuse if the session has already ended.
- Send
cancelto the old node. - Reset the job to ready, clear the owner but keep the epoch, and record where playback should resume.
- If the old node's VM heartbeat is less than 30 seconds old, so the node is still alive, wait up to 8 seconds for it to let go, polling every 500 ms.
- Dispatch to any node except the old one.
Step 3 is written the way it is because of two separate mistakes. Resetting the epoch once broke fencing, since an old encoder's token became valid again. And clearing last_heartbeat_at on failover caused a failover storm: the replacement looked stale the moment it claimed, so it got failed over too, and so on.
Moving onto Kubernetes with KEDA
Update, October 2026: this move was later dropped. The encoder fleet now grows by adding rented VPS hosts through a self-service flow in the admin: the backend prechecks each host over SSH, sizes its encoders by CPU and deploys them. The ownership and failover design above is unchanged. The plan below is kept as it was written.
Today the encoders run on a fleet of cloud VMs with a fixed number of slots each, and the fleet is full. The move to Kubernetes is prepared in the repo as manifests. At the time of writing the cluster itself hasn't been provisioned yet, and the one piece already live in production is the demand endpoint KEDA will read.

Three decisions shaped the setup:
Scale on demand the backend reports, not on a queue. KEDA's Redis scaler would be the obvious choice, but KEDA runs in a separate network and can't reach the private Redis or Postgres. So the backend exposes pendingLive, the number of lives waiting for an encoder, over HTTPS, and KEDA polls it with the metrics-api trigger: minReplicaCount: 0, maxReplicaCount: 50, polling every 15 seconds, target 1.
Throttle scale-up to 2 pods a minute. HPA's default is to double the pods or add four every 15 seconds. Every new encoder pod pulls its whole media file from object storage before it can go live, and across 290 cached files the median was 0.89 GB, the p90 1.73 GB and the largest 8.03 GB. The storage tier was already the bottleneck. On 2026-09-21, encoders re-downloading in lockstep during a failover storm filled node disks, with 64 partial files on one node in two minutes. So scale-up is capped with selectPolicy: Min, so a future policy can't quietly widen it.
Make scale-down pick idle pods. Each encoder patches its own controller.kubernetes.io/pod-deletion-cost annotation, 1,000,000 while it's streaming and 0 when it's idle, every 30 seconds. Scale-down is limited to one pod a minute after a 5-minute window. If a busy pod is removed anyway, it gets 600 seconds of grace: it stops ffmpeg and emits a handoff event with the current playback position, and the backend re-dispatches the live with a seek.
There's one question I still want answered against a real cluster before calling the autoscaling done. pendingLive counts only lives that haven't been claimed yet, so once every live is claimed the metric drops back toward zero, while the pods are still busy. Deletion cost and the handoff path protect the streams if HPA does scale in, but the right fix is probably a demand number that counts running lives too.
Uploads up to 10 GB that survive a bad connection
Customers upload the media from the browser straight to object storage, and some of them are on slow, flaky links. The upload path is S3-style multipart with presigned URLs:
- Session. The backend creates a pending media row. Small files get one presigned PUT. Anything larger starts a multipart upload, and every part URL is presigned up front with a 12-hour expiry, because a 10 GB file on a slow link takes far longer than a 15-minute URL lasts.
- Part size: 6 MB. It used to be 50 MB, sized for a CDN limit that no longer sits on the upload path. At the throughput customers reported, a failed 50 MB part was about 80 seconds of work to redo, and a 6 MB one about 10. A 10 GB file is 1,707 parts, well under the 10,000 cap.
- Resume. The browser stores
{mediaId, uploadId}inlocalStorageunder a fingerprint of name, size and last-modified time, for 24 hours. On retry the backend lists the parts storage already has and re-presigns the rest, and the client uploads only what's missing:
const saved = loadResume(fingerprint(file));
if (saved) {
try {
const { uploadedParts, urls } = await api.resume(saved.uploadId);
await uploadParts(file, urls, { skip: uploadedParts });
return api.complete(saved.mediaId);
} catch (e) {
if (isGone(e)) clearResume(fingerprint(file)); // expired: start fresh
else throw e;
}
}- Transfer. Four parts in flight, three retries each with jittered exponential backoff from 500 ms, and a 503 honours
Retry-After. A 30-second stall watchdog catches uploads that stop moving without erroring. Storage must exposeETagthrough CORS, or the client can't complete the upload. - Cleanup. A bucket lifecycle rule aborts incomplete multipart uploads after two days.
One trap: changing the part size breaks every upload already in flight, because resume works out part offsets from the current size. Roll it out when nothing is mid-upload.
What I'd do again
- One process per stream. It makes ownership, metrics and failure obvious.
- Fence with an epoch, and check it in the worker before and during the side effect, not just at claim time.
- Escalate SIGTERM to SIGKILL. A process blocked on a dead socket will outlive your lease.
- Never reset the epoch or the heartbeat on failover. Both mistakes caused real outages.
- Guard the detector. Two consecutive stale sweeps, a database-degraded grace, and a fleet-wide sanity check stop a slow database from failing over everyone.
- Throttle scale-up to what storage can feed, and use pod deletion cost so scale-down removes idle pods.
- Small multipart parts, URLs that outlive the upload, resume via list-parts, and a lifecycle rule to clean up.
The wider platform, including the SMM order engine and shared SSO, is in the Wavesgroup case study.
Related posts


