What BentoML does
BentoML is an AI inference infrastructure company best known for BentoML, an open-source Python framework for building and serving machine learning and LLM inference workloads with production-oriented reliability. BentoML packages models (“BentOs”) and provides an API/server abstraction plus operational features designed to make inference deployments more predictable across environments and hardware.
More
Commercially, BentoML also offers BentoCloud, an inference management and compute orchestration platform built on top of the open-source serving engine. BentoCloud is designed to let teams deploy inference APIs, batch inference jobs, and multi-component AI systems in their cloud or via “Bring Your Own Cloud” (BYOC) setups, while integrating common inference runtimes (for example, vLLM, TensorRT, and Triton). Strategically, BentoML’s position shifted in 2026 when BentoML joined Modular via a strategic acquisition intended to unify production inference deployment with Modular’s broader AI compute stack. BentoML’s site and Modular’s announcement both frame the move as an acceleration of end-to-end inference platform capabilities, while BentoML remains open source (Apache 2.0).