Project

go-gte

active

Pure Go GTE-Small text embeddings with hand-written SIMD kernels.

Overview

Forked from antirez/gte-pure-C. A pure Go implementation of the GTE-Small text embedding model. Produces 384-dimensional, L2-normalized embeddings suitable for similarity search and clustering.

A single static binary with little allocation in the inference path. All matrix operations use hand-written SIMD assembly (AVX2+FMA on amd64, NEON on arm64) — no gonum, no goroutine churn, no CGo in the default build.

Motivation

This followed asterisk and gave me a reason to work with Go assembly. Non-GPU performance still falls short of what I wanted. Comparing SIMD and NVIDIA execution taught me enough to start go-pherence, even though my potato RTX3060 provides barely any speedup.

I'm pretty happy with it, and it was a great learning experience even if I haven't yet sorted out the vector database I will be using it with.

How it works

Like Salvatore's original, it loads a converted GTE-Small model at startup, tokenises input with a bundled WordPiece tokeniser, runs matrix multiplications (through hand-tuned SIMD kernels), and returns mean-pooled L2-normalized embeddings as a Go slice. Profiling led me away from the gonum BLAS implementation AI had suggested and towards kernels that avoid goroutine and allocation overhead.

Features
🔢
384-dim embeddings

Returns L2-normalized 384-dimensional vectors compatible with GTE-Small's training distribution — plug directly into cosine similarity, FAISS, or any vector store.

⚡
Hand-written SIMD

AVX2+FMA on amd64, NEON on arm64 — non-temporal matmul, fused dot products, and packed transpose kernels written in Go assembly. No CGo required.

🧊
1 allocation per embed

The entire inference path allocates once (uppercase→lowercase token lowering). For all-lowercase input: 0 allocations.

📉
Flat latency

Avoids starting goroutines in the inference path. Batching can further reduce latency variation.

📦
Static binary

Default make produces a fully self-contained static binary with no C dependencies — portable to any amd64 or arm64 target.

Architecture
Input text query or document GGUF weights 23 MB model file Embed pipeline single static binary Tokenizer WordPiece 6L Attention BERT encoder Mean pool L2 normalize 384-dim vector similarity · clustering Pure Go GTE-small — single binary, 1 alloc per embed, flat latency
Posts