Troubleshooting
Start with Manager readiness, then follow the job from durable state to Worker ownership and media execution. Avoid restarting services until you know which boundary failed; retries and reconciliation are designed to preserve evidence.
Manager is live but not ready
curl --include http://localhost:4400/api/health/live
curl --include http://localhost:4400/api/health/ready
- If
databaseis down, verifyDATABASE_URL, TLS policy, migrations, and PostgreSQL reachability. - If
natsis down, verifyNATS_SERVERS, the Manager JWT/NKey pair, the account preload, JetStream availability, and network policy. Token-only NATS cannot authenticate enrolled Workers. - Inspect Manager structured logs for the same request and trace ID.
Jobs remain queued
- List Workers with
connectivity=ONLINE. - Confirm each Worker has a current
READYorDEGRADEDcertification and at least one passing execution path for the plan. - Verify the selected Worker has its durable JetStream consumer and a current NATS lease.
- Check JetStream streams/consumers and outbox publication logs.
- Confirm the Worker is not
DRAININGand has free capacity. If a dispatch times out or a running Worker disappears, inspect the attempt'sDISPATCH_TIMEOUTorWORKER_LOSTfailure code and the next queued attempt.
Legacy H.264 backend names such as libx264, h264_nvenc, and h264_qsv are normalized to portable H.264 intent. The certified-path scheduler still ranks NVIDIA → Intel → CPU and keeps non-codec feature requirements strict.
Worker becomes suspected or offline
Confirm heartbeats reach encode-flow.workers.{workerId}.heartbeat and that the Worker process still uses its current instanceId. Compare network latency with the suspected/offline thresholds. Do not manually reassign an active attempt only because the Worker disconnected; reconciliation and orphan timing prevent duplicate ownership.
Progress disappears after Manager restart
Detailed progress and Worker statistics are in memory. Wait for the next Worker update or inspect durable job/attempt fields. A completed lifecycle remains in PostgreSQL even when the last live FPS, bitrate, or ETA sample is gone.
S3 or FFmpeg failure
- Verify every input, destination, watermark, subtitle, and manifest location uses
s3://bucket/key. - For MinIO, set
S3_ENDPOINTand keepS3_FORCE_PATH_STYLE=true. - Confirm the Worker has disk space in both state and work directories.
- Match
failureCodewith the failing stage and inspectfailureMessageorlogsUri. - Reproduce capability discovery on the same Worker image; package-level FFmpeg differences affect eligibility.
Admin does not update live
Open the Admin server's same-origin GET /api/events endpoint in the browser network panel. A redirect or 401 indicates an expired OIDC session; a 403 from the Manager indicates the token lacks the configured operations Admin scope. Verify reverse proxies disable buffering for SSE. The Admin falls back to ten-second polling while disconnected.
If health always appears unavailable, confirm the Admin proxy can reach the Manager's /api/health/ready path and that CONTROL_PLANE_URL contains the Manager origin without an /api suffix.
Invalid request response
Read every item in details; nested fields use dot paths. Fix the request rather than retrying unchanged. Common failures include invalid UUIDs, unsupported enum values, state combined with states, out-of-range pagination, overlapping removed segments, and more than one burned-in subtitle.