In this field report
How We Test
Benchmark Command documents local AI performance across NVIDIA, AMD, and Intel hardware. The goal is reproducibility: a reader should be able to identify the machine, software stack, workload, measurement and limitations behind every published result.
What every benchmark should disclose
- Machine hardware, operating system and build.
- GPU driver, compute backend and application version.
- Model artifact, quantization, context and workload settings.
- Warm-up policy, sample count and the precise timing definition.
- Power limits, memory pressure and relevant thermal conditions.
Comparison rules
Results are treated as directly comparable only when the model, artifact, backend configuration and workload are materially equivalent. Otherwise they are reported as separate observations, not a winner-and-loser table. Cold starts, warm runs and failed runs are identified rather than silently combined.
Evidence and revisions
Raw logs, structured results, software versions and run receipts are retained for review. Articles should distinguish measured findings from interpretation, disclose anomalies, and record material corrections when drivers, runtimes or test methods change.
Limitations
This is a working homelab test fleet, not a controlled certification laboratory. Driver changes, platform maturity and background system activity can affect results. Repeated runs reduce uncertainty but do not eliminate it.
Publication status
Only linked, completed articles should be treated as published findings. Cards marked as planned describe work in the testing queue; their numbers are not final until the supporting run record has been reviewed.
Raw data
Downloadable result tables and machine-readable run records will be added as each benchmark package is reviewed for publication.
