← Founder Ops

Paper worker

Runs every ACTIVE paper deployment on live Binance data. Without a fresh heartbeat, deploys save but no trades run. See also P0 checklist.

Current setup: Google Cloud Run Jobs + Cloud Scheduler

Since 2026-09-09. Railway's paper-worker hit a persistent platform-side build failure (confirmed not a config, billing, or GitHub-connection issue, Railway's own diagnosis called it an "Infrastructure Error"), so the worker moved to GCP. Same code, same Supabase database, just a different, cheaper, more reliable host.

  1. Cloud Run Job paper-worker (project zengtrade, region us-central1) runs python worker.py --once: one cycle, then exits.
  2. Cloud Scheduler job paper-worker-cycle triggers a fresh execution every 5 minutes.
  3. DATABASE_URL lives in Secret Manager (encrypted, IAM-scoped), not a plain environment variable.
  4. Cost: scale-to-zero between runs, well inside Cloud Run's free tier. Expect $0/month.

Diagnose

No secrets printed by any of these:

gcloud run jobs executions list --job=paper-worker --region=us-central1 --project=zengtrade --limit=10
gcloud scheduler jobs describe paper-worker-cycle --location=us-central1 --project=zengtrade
gcloud logging read 'resource.type="cloud_run_job" AND resource.labels.job_name="paper-worker"' --project=zengtrade --limit=50

Expected when healthy: an execution roughly every 5 minutes, each logging startup heartbeat ok; heartbeat below under 12 minutes old.

If it stops updating

  1. Consecutive FAILED executions in the list above → check the logs command's output. Usual cause: wrong DB password (same symptoms as always, Postgres auth error) or wrong pooler port.
  2. No executions appearing every 5 minutes at all → confirm the Scheduler job is ENABLED and its target URI uses the v2 API path (https://us-central1-run.googleapis.com/v2/projects/zengtrade/locations/us-central1/jobs/paper-worker:run). The older v1/namespaces/.../jobs/...:run path accepts the request but silently never triggers a real execution, cost real debugging time once already.
  3. Rotate the DB password (Supabase → Database settings → reset → copy the session pooler URI, port 5432, not transaction 6543):
read -s -p "Paste the new DATABASE_URL: " DBURL && echo
printf "%s" "$DBURL" | gcloud secrets versions add zengtrade-worker-db-url --data-file=- --project=zengtrade

No redeploy needed: every execution reads the :latest secret version fresh.

Rebuild after a worker code change (Docker isn't required locally, Cloud Build's builds.create API blocked this project as a new-account anti-fraud check; Cloud Shell's pre-installed Docker sidesteps it):

cd Zengtrade-V2/saas/worker   # git pull first if already cloned
docker build -t us-central1-docker.pkg.dev/zengtrade/zengtrade/paper-worker:latest .
docker push us-central1-docker.pkg.dev/zengtrade/zengtrade/paper-worker:latest

If docker push fails with connect: connection refused to a googleapis.com IP, just retry: a transient Google network blip, not a config problem. The Cloud Run Job always pulls :latest on its next scheduled execution, no separate deploy step.

Manual test run

Run one cycle immediately, without waiting for the next 5-minute tick:

gcloud run jobs execute paper-worker --region=us-central1 --project=zengtrade --wait

Stop / pause (safe)

Does not delete deployments or historical trades: users simply stop getting new paper fills until resumed.

gcloud scheduler jobs pause paper-worker-cycle --location=us-central1 --project=zengtrade   # stop
gcloud scheduler jobs resume paper-worker-cycle --location=us-central1 --project=zengtrade  # resume

Legacy: Railway (superseded 2026-09-09)

The original paper-worker service in this Railway project is no longer the source of truth: its builds got permanently stuck on Railway's side. Left in place rather than deleted (costs nothing while failing to build); safe to remove whenever convenient. saas/worker/Dockerfile and railway.toml still describe how to redeploy there if ever needed as a fallback.

Meanwhile: parallel work

Independent of worker status:

Verify

  1. Heartbeat below should show < 12 minutes
  2. /ops → Paper worker gate = ✓
  3. Test: signup → deploy → wait ~5 min → trades
  4. /admin → Worker tile = Live
  5. Full runbook: docs/WORKER_RECOVERY.md

Checking worker heartbeat…