Methodology

Publication standardBy Benchmark CommandPublished About 2 min read
In this field report

How We Test

Benchmark Command documents local AI performance across NVIDIA, AMD, and Intel hardware. The goal is reproducibility: a reader should be able to identify the machine, software stack, workload, measurement and limitations behind every published result.

What every benchmark should disclose

  • Machine hardware, operating system and build.
  • GPU driver, compute backend and application version.
  • Model artifact, quantization, context and workload settings.
  • Warm-up policy, sample count and the precise timing definition.
  • Power limits, memory pressure and relevant thermal conditions.

Comparison rules

Results are treated as directly comparable only when the model, artifact, backend configuration and workload are materially equivalent. Otherwise they are reported as separate observations, not a winner-and-loser table. Cold starts, warm runs and failed runs are identified rather than silently combined.

Evidence and revisions

Raw logs, structured results, software versions and run receipts are retained for review. Articles should distinguish measured findings from interpretation, disclose anomalies, and record material corrections when drivers, runtimes or test methods change.

Limitations

This is a working homelab test fleet, not a controlled certification laboratory. Driver changes, platform maturity and background system activity can affect results. Repeated runs reduce uncertainty but do not eliminate it.

Publication status

Only linked, completed articles should be treated as published findings. Cards marked as planned describe work in the testing queue; their numbers are not final until the supporting run record has been reviewed.

Raw data

Downloadable result tables and machine-readable run records will be added as each benchmark package is reviewed for publication.