Standards consortium · Funders · Dover, Delaware · Est. 2020
MLCommons is an open, community-driven engineering consortium focused on making AI better through neutral benchmarking, large-scale open datasets, and supporting engineering standards for how those benchmarks and datasets are packaged and evaluated. Its mission is to accelerate AI innovation and increase positive societal impact by democratizing ML via open industry-standard benchmarks for measuring quality and performance and by building open, large-scale, diverse datasets intended to improve AI models.
The organization frames its work around accelerating industry-wide adoption of measurable, reproducible AI evaluation and includes both performance benchmarking and AI risk & reliability efforts. Institutionally, MLCommons functions as a neutral coordinator: it convenes a diverse set of “founding Members and Affiliates” (including startups, leading companies, academics, and non-profits) to co-develop benchmark suites and related tooling. MLCommons emphasizes open collaboration and published policies governing benchmark operations and participation. Technically, MLCommons’ differentiation is the combination of (1) benchmark suites spanning multiple application tiers and modalities (e.g., training, inference, mobile, endpoints, storage, and medical benchmarking), (2) dataset and metadata/data-tooling standards intended to improve discoverability and reproducibility, and (3) reliability and safety-oriented benchmarking initiatives. Recent releases continue to expand benchmark scope (e.g., MLPerf Storage v3.0 adding KV-cache and vector-database tests and an S3 access layer) and to add evaluation approaches intended to improve trustworthiness (e.g., discussion and releases related to secrecy-by-design and reliability evaluation methods).
No indexed openings right now. Check the careers page ↗.
Benchmark suite now covers the full range of AI workloads for storage systems The post MLCommons Releases New MLPerf Storage v3.0 Benchmark Results appeared first on MLCommons .
The key to trustworthy AI evaluation is secrecy by designAILuminate’s critical role in the first double-blind reliability evaluation of a proprietary AI model The post The key to trustworthy AI evaluation is secrecy by design appeared first on MLCommons .
Introducing the MLPerf End-to-End RAG Inference BenchmarkBenchmarking Retrieval-Augmented Generation End to End — from Building the Vector Database to Serving an Iterative, Multi-Hop Question-Answering Inference Pipeline The post Introducing the MLPerf End-to-End RAG Inference Benchmark appeared first on MLCommons .
MLPerf Client v2.0 Expands AI PC Benchmarking with Image Generation and Agentic AINew benchmark categories for generative AI and agentic workflows arrive alongside updated LLM tests, as the industry standard for measuring PC AI performance evolves. The post MLPerf Client v2.0 Expands AI PC Benchmarking with Image Generation and Agentic AI appeared first on MLC
How to Tell When a Benchmark Is Worth TrustingAn enterprise guide from the people who build them The post How to Tell When a Benchmark Is Worth Trusting appeared first on MLCommons .
MLPerf Endpoints v0.7: A Foundation ReleaseFrom demonstration at GTC to foundation release: new results, automated submission pipelines, and a roadmap to v1.0 with rolling submissions later this year. The post MLPerf Endpoints v0.7: A Foundation Release appeared first on MLCommons .
MedPerf Meets Google Cloud Confidential Computing: Secure AI Benchmarking for Brain Tumor ResearchAt Google Cloud Next 2026 in Las Vegas, MLCommons and Google Cloud demonstrated a powerful new capability for trustworthy medical AI - one that protects patient data, model IP, and benchmark integrity all at once. The post MedPerf Meets Google Cloud Confidential Computing: Secure
Call for Submission: Edge Agentic Inference Benchmark for MLPerf Inference v6.1Benchmarking Multi-Turn Agentic LLMs on a Single Edge Accelerator The post Call for Submission: Edge Agentic Inference Benchmark for MLPerf Inference v6.1 appeared first on MLCommons .
Agentic Inference for MLPerf InferenceA new multi-turn benchmark for measuring LLM serving systems under growing context, and closed-loop agent workflows. The post Agentic Inference for MLPerf Inference appeared first on MLCommons .
The Benchmark Behind the Next Wave of Ultra-Low-Power AIMLPerf Tiny: Benchmarking AI at the Edge The post The Benchmark Behind the Next Wave of Ultra-Low-Power AI appeared first on MLCommons .
MLCommons Releases MLPerf Training v6.0 ResultsNew benchmarks and increased diversity of submissions reflect important changes in AI ecosystem The post MLCommons Releases MLPerf Training v6.0 Results appeared first on MLCommons .
MLCommons Releases MLPerf Mobile v6.0 with New Generative AI Benchmarks for On-Device LLMsTest LLM inference natively on mobile devices with new standardized benchmarks and expanded NPU acceleration. The post MLCommons Releases MLPerf Mobile v6.0 with New Generative AI Benchmarks for On-Device LLMs appeared first on MLCommons .
MLCommons announced MLPerf Storage v3.0 results, describing new tests (KV cache and vector database) and added support for an S3 object storage access layer, along with new workload coverage for storage patterns relevant to AI systems.
Introducing the MLPerf End-to-End RAG Inference BenchmarkMLCommons introduced an end-to-end RAG inference benchmark, aiming to standardize RAG evaluation across retrieval and generation components in a single measurement framework.
MLPerf Client v2.0 Expands AI PC Benchmarking with Image Generation and Agentic AIMLCommons released MLPerf Client v2.0, describing expanded coverage for AI PC benchmarking that includes image generation and agentic AI scenarios.
How to Tell When a Benchmark Is Worth TrustingMLCommons published guidance on benchmark trustworthiness, focusing on evaluation integrity and what to look for when using benchmark results to make decisions.
AILuminate and The First Double-Blind Reliability Evaluation of a Proprietary AI ModelMLCommons published a reliability-evaluation update describing double-blind evaluation context and its role in AILuminate-style evaluation, emphasizing evaluation integrity and trustworthiness.
MLPerf Endpoints v0.7: A Foundation ReleaseMLCommons released MLPerf Endpoints v0.7, described as a foundation release for the endpoints benchmark line.