← All stories
● Covered by 1 source · 1 reportLow impact1 neutral

Janus: A Go Binary for Running GGUF Models with Vulkan on AMD/Intel/Nvidia GPUs

🔄 Updated 16h ago
New to BrevFeed? We gather this story from every outlet covering it into one summary — ranked by real-world impact, not just the latest headline — so you never miss what matters. What is BrevFeed? →

Key points

  • Single Go binary for GGUF model inference.
  • Utilizes Vulkan for GPU acceleration (AMD, Intel, Nvidia).
  • Offers OpenAI-compatible API and built-in web UI.
  • Includes tools for file operations, commands, and OCR.

Local AI Inference with Go

Janus is a recently released Go binary designed to run .gguf models directly on a user's machine. It supports GPU acceleration via Vulkan, compatible with AMD, Intel, and Nvidia graphics cards, and includes a CPU fallback option. This approach aims to provide a self-contained environment for local AI inference, removing the need for Python, Docker, or Ollama, though Ollama can be used as a backend if preferred.

OpenAI-Compatible API and Web UI

The tool exposes an OpenAI-compatible API, including endpoints like `/v1/chat/completions` and `/v1/models`, which allows integration with various OpenAI clients such as Cursor or Cline. Additionally, Janus features a built-in web UI that offers functionalities like an Assistant, Chat, Kernel (for tool loops), configuration settings, memory management, and skill definitions. This dual interface supports different user workflows for interacting with local models.

Integrated Tools and Model Management

Janus incorporates a suite of built-in tools for various tasks, including reading and writing files, executing commands, performing mathematical operations, handling DOCX input/output, generating PDF output, and OCR capabilities via Tesseract. Users can hot-swap .gguf models through the UI without restarting the application. Optional basic authentication is available for administrative endpoints.

Setup and Configuration

Installation involves cloning the Git repository, building the executable, and placing .gguf model files in a designated folder. A model downloader is provided for convenience. Configuration is managed through an `.env` file, where users specify the inference backend (e.g., `vulkan` or `cpu`), the model path, and other parameters like `JANUS_MAX_TOKENS` and `JANUS_AUTH`. The first startup loads the model into VRAM, which can take 10-60 seconds depending on model size and disk speed.

✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →

The daily brief

One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.

One email a day. Unsubscribe in one click, any time.

Today's brief

Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.

~34 min · 27 stories · Oct 02

▶ Play today's brief Listen on Spotify

New every morning, and the back catalogue is archived by date.

Reporting from

Janus is a new single Go binary that enables local execution of GGUF models using Vulkan for GPU acceleration across AMD, Intel, and Nvidia hardware, or CPU fallback. It provides an OpenAI-compatible API and a web UI, offering a self-contained solution for local AI inference without external dependencies like Python or Docker.