Project

go-pherence

active

Local inference in Go — LLMs, embeddings, speech and experimental vision models on CPU and GPU.

Overview

go-pherence runs language, embedding and speech models on local hardware, with experimental support for vision and image generation. It provides Go libraries, command-line tools and an OpenAI-compatible server, including LLaMA, Qwen and Gemma decoders, BERT/GTE embeddings, Whisper transcription and MOSS transcription with speaker labels.

CPU execution uses Go and hand-written assembly. NVIDIA support loads PTX through the installed driver without cgo or a CUDA toolkit. Vulkan and embedded accelerator support are opt-in and model-dependent.

How it works

Loaders read model configuration, tokenisers and weights from safetensors, selected GGUF layouts and associated files. Model packages implement attention, generation loops, speech processing and image schedulers, sharing kernels and memory management rather than forcing every model through one graph engine.

CPU kernels use AVX2, NEON or RISC-V vector instructions where supported, with scalar Go fallbacks. NVIDIA execution keeps weights and intermediate state on the GPU where the model implementation allows it. Quantised weights retain their packed layout where supported, reducing memory traffic without first expanding the whole checkpoint.

The speech tools can produce transcripts and subtitles with timestamps and speaker labels. Their normal media input uses FFmpeg; library callers can instead select a pure-Go go-264 adapter for PCM WAV and a limited AAC-LC MP4/M4A subset.

Features
🦙
Local LLMs

Dense and mixture-of-experts LLaMA, Qwen and Gemma-family models, with command-line generation, interactive chat and an OpenAI-compatible server.

📦
Checkpoint formats

MLX and GPTQ 4-bit weights, BF16/F16/F32 safetensors and selected GGUF layouts. Supported formats depend on the model architecture.

⚡
CPU and NVIDIA execution

AVX2, NEON and RVV kernels with scalar fallbacks; runtime-loaded PTX for supported NVIDIA operations. No Python inference runtime required.

🧠
Embeddings and extraction

BERT/GTE embeddings, including the GTE-small model used by go-gte, and GLiNER 2.5 entity, classification, relation and record extraction.

🎙
Speech and speaker labels

Whisper transcription or translation to WebVTT, with optional diarisation. MOSS provides native transcription with timestamps and speaker labels local to each recording.

🎨
Experimental generation

Ideogram 4 has a native CPU/SIMD image-generation implementation. DiffusionGemma has a partial native text-generation implementation; neither offers complete upstream model coverage.

🔧
Library use

Model loaders, tensor operations and execution backends can be used from Go applications. Jevlike adds choice scoring, training and reusable visual adapters.

🧪
Model-specific limits

Some families, including Qwen3-TTS and MiniCPM-V/O, have inspection or preprocessing support without full inference. The supported-model guide lists the runnable models and their limits.

Architecture
Model weights safetensors / GGUF Model execution model-specific graphs Model code layers / loops Shared state tensor / KV Compute backends CPU kernels Go + assembly NVIDIA driver / PTX Inference tools Text LLMs / vectors Speech text / speakers Shared loaders and kernels for local model inference
Posts