
ZML ships a free inference server that runs open-source LLMs across five chip families
ZML/LLMD, a free — though not open-source — inference server from the 20-person Paris startup ZML, runs open-source models across Nvidia, AMD, Google TPU, Apple Metal, and Intel Arc from one server. A developer can serve LLMs on a mixed-vendor fleet instead of a single-vendor GPU stack, though ZML disclosed no benchmark numbers to back the peak-performance claim.
Source: techcrunch.com ↗