Joint Model & Hardware Evaluation
Evaluate model quality and hardware performance together to guide edge model and device selection.
Edge AI Evaluation Infrastructure
JishuBench combines unified orchestration with a Host-Target architecture to create reproducible evaluation evidence for edge AI models, devices, and inference stacks.
View on GitHubHost
Orchestration · Datasets · Scoring
BENCHMARKS
Target
Inference · Hardware metrics
MODELS
DEVICES
Connect model quality, hardware performance, and real workloads with data that can be compared and trusted.
Evaluate model quality and hardware performance together to guide edge model and device selection.
Bring leading evaluation frameworks into one workflow for multiple benchmark types.
Run evaluation on an easy-to-deploy Host while the Target only runs model inference and lightweight integration.
The Host orchestrates evaluation while the Target serves inference, connected through an OpenAI-compatible interface.
Install JishuBench, evaluation frameworks, and datasets on a workstation or server for preflight, orchestration, and scoring.
Run llama-server on the edge device to provide inference to the Host, with optional hardware metric collection.
Phase 1 Measured Results
Results are from single task runs on internal project devices and are for reference only.
| Model | mmmu_val | mmbench | realworldqa | pope | textvqa | mmstar | ocrbench | Total time |
|---|---|---|---|---|---|---|---|---|
| 模型 A-35B-A3B | 49.4% | 86.7% | 77.9% | 88.8% | 84.5% | 69.6% | 86.0% | 263 min |
| 模型 B-26B-A4B | 54.4% | 84.4% | 66.4% | 86.8% | 68.4% | 70.0% | 82.7% | 300 min |
| 模型 C-9B | 43.4% | 83.9% | 74.9% | 89.8% | 83.2% | 68.2% | 85.1% | 271 min |
| 模型 D-27B | 52.6% | 87.7% | 76.0% | 89.0% | 84.5% | 72.4% | 85.8% | 570 min |
| 模型 E-31B | 58.9% | 87.0% | 68.9% | 88.4% | 71.7% | 74.4% | 85.5% | 657 min |
Evaluated with lmms-eval. The table shows primary metric scores for mmmu_val, mmbench, realworldqa, pope, textvqa, mmstar, and ocrbench, along with 7-task parallel batch wall time.
| Model | airline | retail | telecom | pass@1 |
|---|---|---|---|---|
| 模型 A-35B-A3B | 74.0% | 80.7% | 90.4% | 83.5% |
| 模型 B-26B-A4B | 44.7% | 78.4% | 13.2% | 44.2% |
Shows pass@1 reward across the airline, retail, and telecom domains.