Treat repeated 503 responses as an availability investigation. Increasing preload concurrency or retries can make recovery harder.
Capture the timeline and affected routes
Record when preload started, which routes returned errors and whether normal visitors were affected. Save response status, relevant headers and host log timestamps. A 503 indicates temporary unavailability but does not identify the underlying cause by itself. Check the host’s incident information and maintenance state before assuming that the preload tool is responsible.
Reduce the suspected background workload
If the timing correlates, pause or reduce preload using the responsible tool’s supported controls. Avoid repeatedly launching a full-site rebuild to test the hypothesis. Verify key public routes and store functions during recovery. Preserve the configuration and logs so a maintainer can compare the workload with error windows rather than relying on memory.
Compare host capacity signals
Ask the host to examine relevant worker, CPU, memory and backend signals for the same interval. Look for slow uncached routes or upstream dependencies that make regeneration expensive. Do not expose full database dumps or customer logs in a performance report. A capacity limit, maintenance operation and failing external integration require different corrections even if they all appear as unavailable responses.
Restart with a bounded route sample
On staging or within host-approved capacity, test a small set of representative public routes at low supported concurrency. Include a heavier route, but exclude customer-specific pages. Record completion time and errors before increasing the scope. Keep a stop condition for availability degradation. A successful tiny test is a starting point for capacity planning, not proof that an unrestricted full-site warmup is safe.
Document the operating policy
Agree who owns preload scheduling, which routes qualify and how it interacts with publication purges. Include a recovery action and an escalation contact. Verify that the chosen schedule does not overlap another expensive job. Restore the earlier stable workload if errors recur. The final change should reduce operational risk while retaining useful warm entries, rather than maximizing a preload counter.
Symptom-to-cause worksheet
| What you observe | What to investigate | Next check |
|---|---|---|
| Errors align with warmup window | Background regeneration may exceed capacity | Pause supported preload and correlate logs |
| Errors persist after workload stops | Independent service issue possible | Check host incident and backend health |
| Only one route triggers failures | Expensive or failing generation path | Investigate that route separately |
Worked investigation scenario
This is an illustrative case, not a measured Velonic result. Suppose errors begin during a full-site warmup and subside after it is paused. This correlation supports investigating capacity, but does not identify the precise bottleneck. Ask the host to compare the same timestamps with worker and backend signals. Restart only a small supported public-route sample and keep a stop condition. If errors persist without warmup, widen the availability investigation instead of repeatedly changing preload settings.
Put it into practice
- Correlate timestamps before assigning cause.
- Reduce workload instead of increasing retries.
- Restart with a small sample and explicit stop condition.
AI-assisted educational content prepared for the Velonic resource library. How these guides are prepared.