Job lifecycle

A job is the user-visible unit of work. An attempt is one ownership window for one Worker. Retries create new attempts so historical execution remains visible and a stale Worker cannot finalize newer work.

Job states

StateMeaning
QUEUEDPersisted and waiting for dispatch.
DISPATCHINGAn attempt is being offered to its selected pool.
RUNNINGA Worker owns the current attempt.
WAITING_FOR_STORAGEOutput transfers are paused; files remain on the original Worker until retention expires.
WAITING_FOR_REVIEWOutput requires a later review decision.
CANCEL_REQUESTEDCancellation is durably requested for active work.
CANCELLEDWork ended because of cancellation.
COMPLETEDThe current attempt reported output artifacts and an uploaded manifest.
FAILEDNo current attempt can complete the job.

Attempt ownership and fencing

Each attempt carries a positive fencingToken. The Worker repeats the job ID, attempt ID, Worker identity, token, and monotonically increasing event sequence on lifecycle events. The Manager rejects messages that do not match current ownership.

This protects the platform when a Worker is partitioned, a JetStream message is redelivered, or an old process resumes after the Manager assigned a retry.

Processing stages

These stages describe the work performed. For incremental jobs, encoding, verification, and uploading can overlap; stage progress and output-unit progress are reported separately.

  1. 01
    Preparing

    Create the isolated workspace and validate the plan.

  2. 02
    Input probe

    Inspect source streams and duration.

  3. 03
    Input transfer

    Copy source objects into the Worker workspace.

  4. 04
    Transcoding

    Run FFmpeg and publish live progress.

  5. 05
    Packaging

    Assemble file or HLS output structures.

  6. 06
    Uploading

    Write artifacts and the manifest to S3.

  7. 07
    Verifying

    Confirm that the execution produced at least one output artifact.

  8. 08
    Finalizing

    Publish the terminal event to the Manager.

  9. 09
    Cleanup

    Remove local media and spool state when safe.

Incremental execution and upload recovery

New portal jobs require incremental-output-recovery-v1. Encoding and uploading overlap: verified output units transfer as they finish, the HLS master follows its dependencies, and the manifest follows all artifacts. A transient encoding failure retries the affected unit once; completed units remain available on that attempt. GPU failures can create a new attempt on another compatible target, which starts encoding again.

Permanent or exhausted upload failures move the job and attempt to WAITING_FOR_STORAGE and release encoding capacity. Retention lasts 24 hours from the first storage pause, subject to the job deadline; further pauses and resume requests do not extend it. After fixing storage, use the resume uploads endpoints. Resume preserves the attempt and destination and waits for the original Worker and capacity. Missing retained files or expired retention require a full retry.

Cancellation

Queued jobs can be cancelled without contacting a Worker. Active jobs transition to CANCEL_REQUESTED; the Manager writes a durable command through the outbox, and the Worker aborts FFmpeg before publishing CANCELLED.

POST /api/v1/operations/jobs/{jobId}/cancel returns 202 Accepted for operators. The customer equivalent is POST /api/v1/projects/{projectId}/jobs/{jobId}/cancel; the final state can be asynchronous.

Reconnect reconciliation

For execution plans without a recovery policy, startup reconciliation reports locally recorded attempts and pending terminal event IDs. The Manager answers CONTINUE, CANCEL, or STALE for each attempt. A non-CONTINUE decision removes the active record, and pending terminal events are replayed. This legacy path does not restart interrupted FFmpeg execution.

Recovery-enabled attempts use a separate polling path. The original Worker replays pending events, requests renewed ownership and storage access, then resumes from retained checkpoints. An interrupted encoding unit restarts; verified completed units and confirmed multipart parts can be reused. Persist both Worker state and work directories across restarts.