Skip to content

Debug trials

Most trial failures come from one of three places: the image cannot start, the CLI arguments do not match, or no valid metrics were printed.

The trial log should answer three questions:

  1. Container start: Did the container start?
  2. Arguments: Did the program receive the expected --hpo-* arguments (or plain flags if configured)?
  3. Metrics: Did stdout include at least one valid hpo.metrics.* line for each objective?
Failure Area Description
unknown argument parser The dashboard parameter slug does not match your CLI parser.
missing metric collector The program finished without printing a valid hpo.metrics.* line.
invalid JSON collector The value after = is not JSON-serializable.
data not found runtime The image cannot access a dataset or configuration file it needs.
image pull / not found registry Tag missing from org Images, or image outside org project on non-Enterprise plans.
no new trials billing / limits Insufficient credits, failure budget, spend limit, or paused experiment.

Run the same command shape on your machine before launching a large experiment.

Terminal window
docker run --rm your-image:latest \
--hpo-lookback-window=50 \
--hpo-risk-multiplier=1.4

Confirm the output contains metric lines:

hpo.metrics.objective=0.84
hpo.metrics.duration_seconds=42.1

For exact collector rules, see Metric format. For a broader checklist, see Troubleshooting.