Troubleshooting

Start with Manager readiness, then follow the job from durable state to Worker ownership and media execution. Avoid restarting services until you know which boundary failed; retries and reconciliation are designed to preserve evidence.

Manager is live but not ready

curl --include http://localhost:4400/api/health/live
curl --include http://localhost:4400/api/health/ready
  • If database is down, verify DATABASE_URL, TLS policy, migrations, and PostgreSQL reachability.
  • If nats is down, verify NATS_SERVERS, the Manager JWT/NKey pair, the account preload, JetStream availability, and network policy. Token-only NATS cannot authenticate enrolled Workers.
  • Inspect Manager structured logs for the same request and trace ID.

Jobs remain queued

  1. List Workers with connectivity=ONLINE.
  2. Confirm each Worker has a current READY or DEGRADED certification and at least one passing execution path for the plan.
  3. Verify the selected Worker has its durable JetStream consumer and a current NATS lease.
  4. Check JetStream streams/consumers and outbox publication logs.
  5. Confirm the Worker is not DRAINING and has free capacity. If a dispatch times out or a running Worker disappears, inspect the attempt's DISPATCH_TIMEOUT or WORKER_LOST failure code and the next queued attempt.

Legacy H.264 backend names such as libx264, h264_nvenc, and h264_qsv are normalized to portable H.264 intent. The certified-path scheduler still ranks NVIDIA → Intel → CPU and keeps non-codec feature requirements strict.

Worker becomes suspected or offline

Confirm heartbeats reach encode-flow.workers.{workerId}.heartbeat and that the Worker process still uses its current instanceId. Compare network latency with the suspected/offline thresholds. Do not manually reassign an active attempt only because the Worker disconnected; reconciliation and orphan timing prevent duplicate ownership.

Progress disappears after Manager restart

Detailed progress and Worker statistics are in memory. Wait for the next Worker update or inspect durable job/attempt fields. A completed lifecycle remains in PostgreSQL even when the last live FPS, bitrate, or ETA sample is gone.

S3 or FFmpeg failure

  • Verify every input, destination, watermark, subtitle, and manifest location uses s3://bucket/key.
  • For MinIO, set S3_ENDPOINT and keep S3_FORCE_PATH_STYLE=true.
  • Confirm the Worker has disk space in both state and work directories.
  • Match failureCode with the failing stage and inspect failureMessage or logsUri.
  • Reproduce capability discovery on the same Worker image; package-level FFmpeg differences affect eligibility.

Admin does not update live

Open the Admin server's same-origin GET /api/events endpoint in the browser network panel. A redirect or 401 indicates an expired OIDC session; a 403 from the Manager indicates the token lacks the configured operations Admin scope. Verify reverse proxies disable buffering for SSE. The Admin falls back to ten-second polling while disconnected.

If health always appears unavailable, confirm the Admin proxy can reach the Manager's /api/health/ready path and that CONTROL_PLANE_URL contains the Manager origin without an /api suffix.

Invalid request response

Read every item in details; nested fields use dot paths. Fix the request rather than retrying unchanged. Common failures include invalid UUIDs, unsupported enum values, state combined with states, out-of-range pagination, overlapping removed segments, and more than one burned-in subtitle.