go-pherence runs language, embedding and speech models on local hardware, with experimental support for vision and image generation. It provides Go libraries, command-line tools and an OpenAI-compatible server, including LLaMA, Qwen and Gemma decoders, BERT/GTE embeddings, Whisper transcription and MOSS transcription with speaker labels.
CPU execution uses Go and hand-written assembly. NVIDIA support loads PTX through the installed driver without cgo or a CUDA toolkit. Vulkan and embedded accelerator support are opt-in and model-dependent.
Loaders read model configuration, tokenisers and weights from safetensors, selected GGUF layouts and associated files. Model packages implement attention, generation loops, speech processing and image schedulers, sharing kernels and memory management rather than forcing every model through one graph engine.
CPU kernels use AVX2, NEON or RISC-V vector instructions where supported, with scalar Go fallbacks. NVIDIA execution keeps weights and intermediate state on the GPU where the model implementation allows it. Quantised weights retain their packed layout where supported, reducing memory traffic without first expanding the whole checkpoint.
The speech tools can produce transcripts and subtitles with timestamps and speaker labels. Their normal media input uses FFmpeg; library callers can instead select a pure-Go go-264 adapter for PCM WAV and a limited AAC-LC MP4/M4A subset.
Dense and mixture-of-experts LLaMA, Qwen and Gemma-family models, with command-line generation, interactive chat and an OpenAI-compatible server.
MLX and GPTQ 4-bit weights, BF16/F16/F32 safetensors and selected GGUF layouts. Supported formats depend on the model architecture.
AVX2, NEON and RVV kernels with scalar fallbacks; runtime-loaded PTX for supported NVIDIA operations. No Python inference runtime required.
BERT/GTE embeddings, including the GTE-small model used by go-gte, and GLiNER 2.5 entity, classification, relation and record extraction.
Whisper transcription or translation to WebVTT, with optional diarisation. MOSS provides native transcription with timestamps and speaker labels local to each recording.
Ideogram 4 has a native CPU/SIMD image-generation implementation. DiffusionGemma has a partial native text-generation implementation; neither offers complete upstream model coverage.
Model loaders, tensor operations and execution backends can be used from Go applications. Jevlike adds choice scoring, training and reusable visual adapters.
Some families, including Qwen3-TTS and MiniCPM-V/O, have inspection or preprocessing support without full inference. The supported-model guide lists the runnable models and their limits.