Job lifecycle
A job is the user-visible unit of work. An attempt is one ownership window for one Worker. Retries create new attempts so historical execution remains visible and a stale Worker cannot finalize newer work.
Job states
| State | Meaning |
|---|---|
QUEUED | Persisted and waiting for dispatch. |
DISPATCHING | An attempt is being offered to its selected pool. |
RUNNING | A Worker owns the current attempt. |
WAITING_FOR_STORAGE | Output transfers are paused; files remain on the original Worker until retention expires. |
WAITING_FOR_REVIEW | Output requires a later review decision. |
CANCEL_REQUESTED | Cancellation is durably requested for active work. |
CANCELLED | Work ended because of cancellation. |
COMPLETED | The current attempt reported output artifacts and an uploaded manifest. |
FAILED | No current attempt can complete the job. |
Attempt ownership and fencing
Each attempt carries a positive fencingToken. The Worker repeats the job ID, attempt ID, Worker identity, token, and monotonically increasing event sequence on lifecycle events. The Manager rejects messages that do not match current ownership.
This protects the platform when a Worker is partitioned, a JetStream message is redelivered, or an old process resumes after the Manager assigned a retry.
Processing stages
These stages describe the work performed. For incremental jobs, encoding, verification, and uploading can overlap; stage progress and output-unit progress are reported separately.
- 01Preparing
Create the isolated workspace and validate the plan.
- 02Input probe
Inspect source streams and duration.
- 03Input transfer
Copy source objects into the Worker workspace.
- 04Transcoding
Run FFmpeg and publish live progress.
- 05Packaging
Assemble file or HLS output structures.
- 06Uploading
Write artifacts and the manifest to S3.
- 07Verifying
Confirm that the execution produced at least one output artifact.
- 08Finalizing
Publish the terminal event to the Manager.
- 09Cleanup
Remove local media and spool state when safe.
Incremental execution and upload recovery
New portal jobs require incremental-output-recovery-v1. Encoding and uploading overlap: verified output units transfer as they finish, the HLS master follows its dependencies, and the manifest follows all artifacts. A transient encoding failure retries the affected unit once; completed units remain available on that attempt. GPU failures can create a new attempt on another compatible target, which starts encoding again.
Permanent or exhausted upload failures move the job and attempt to WAITING_FOR_STORAGE and release encoding capacity. Retention lasts 24 hours from the first storage pause, subject to the job deadline; further pauses and resume requests do not extend it. After fixing storage, use the resume uploads endpoints. Resume preserves the attempt and destination and waits for the original Worker and capacity. Missing retained files or expired retention require a full retry.
Cancellation
Queued jobs can be cancelled without contacting a Worker. Active jobs transition to CANCEL_REQUESTED; the Manager writes a durable command through the outbox, and the Worker aborts FFmpeg before publishing CANCELLED.
POST /api/v1/operations/jobs/{jobId}/cancel returns 202 Accepted for operators. The customer equivalent is POST /api/v1/projects/{projectId}/jobs/{jobId}/cancel; the final state can be asynchronous.
Reconnect reconciliation
For execution plans without a recovery policy, startup reconciliation reports locally recorded attempts and pending terminal event IDs. The Manager answers CONTINUE, CANCEL, or STALE for each attempt. A non-CONTINUE decision removes the active record, and pending terminal events are replayed. This legacy path does not restart interrupted FFmpeg execution.
Recovery-enabled attempts use a separate polling path. The original Worker replays pending events, requests renewed ownership and storage access, then resumes from retained checkpoints. An interrupted encoding unit restarts; verified completed units and confirmed multipart parts can be reused. Persist both Worker state and work directories across restarts.