← All stories
● Covered by 1 source · 1 reportLow impact1 neutral

Hetzner experiments with LLM inference service using OpenAI-compatible API

🔄 Updated 1h ago
New to BrevFeed? We gather this story from every outlet covering it into one summary — ranked by real-world impact, not just the latest headline — so you never miss what matters. What is BrevFeed? →

Key points

  • Hetzner launched an experimental LLM inference service.
  • The service uses an OpenAI-compatible API.
  • Currently, only the Qwen/Qwen3.6-35B-A3B-FP8 model is available.
  • The experiment lacks billing, SLA, or production guarantees.

Hetzner's LLM Inference Experiment

Hetzner has introduced an experimental service for Large Language Model (LLM) inference. This offering is not a finished product but an early release designed to collect user feedback and assess technical requirements. The company aims to understand demand, scaling behavior, necessary features, and load capacity.

OpenAI-Compatible API

The experimental service provides an OpenAI-compatible API, allowing users to integrate it with existing OpenAI clients by changing the base URL and using an API token. This approach simplifies testing for developers familiar with OpenAI's ecosystem.

Available Model and Features

Currently, the only model available is Qwen/Qwen3.6-35B-A3B-FP8, a 35-billion-parameter Mixture-of-Experts model with 3 billion active parameters. It supports text and image input, features a 262K context window, and uses FP8-quantized weights. This model size is suitable for experimental purposes, balancing utility with resource requirements.

Experimental Status and Limitations

Hetzner emphasizes that this is an experiment, meaning there are no billing mechanisms, Service Level Agreements (SLAs), or production guarantees in place. Users are advised against deploying production AI workloads on this service due to its early-stage nature. A tutorial for connecting OpenCode to the API is also available for users who wish to test without writing code.

✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →

The daily brief

One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.

One email a day. Unsubscribe in one click, any time.

Today's brief

Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.

~15 min · 15 stories · Jul 24

▶ Play today's brief Listen on Spotify

New every morning, and the back catalogue is archived by date.

Reporting from

Hetzner is testing an experimental LLM inference service, providing an OpenAI-compatible API for users to interact with a single model, Qwen/Qwen3.6-35B-A3B-FP8. This initiative aims to gather data on user interest, system scalability, and feature requirements before a formal product launch.