Best local embedding model for job-posting search on a 16 GB M4 Mac?

Hi all,

I’m building a local job-search assistant and I’m looking for the best
embedding model for it. Everything runs on my own machine, so speed and
memory matter as much as quality.

The setup

  • Corpus: about 10,600 English job postings, split into 500-character
    chunks: 141,603 chunks, about 13.6M tokens (roughly 96 tokens per chunk).
  • Queries: short questions like “backend roles using Kafka”, plus
    paraphrased ones that describe the work without the posting’s own words.
  • Search: Postgres + pgvector. Keywords pick the candidates and the
    embedding model orders them, by each posting’s best-matching chunk.
  • Hardware: Apple M4, 16 GB RAM (MPS). No cloud, no paid APIs.

What I measured (56 postings, two questions each, pool of 400 postings;
“top 5” = the posting the question was written from is in the top 5)

Model Natural top 5 Paraphrased top 5 Speed
bge-small-en-v1.5 (current) 77% 32% 366 chunks/s
bge-large-en-v1.5 79% 45% 34 chunks/s
Qwen3-Embedding-4B (Q4 GGUF) 91% 71% 3.3 chunks/s

Qwen3-Embedding-4B is clearly the best, but it would take about 12 hours
to embed my corpus, against about 6 minutes for bge-small. I’m testing
Qwen3-Embedding-0.6B now (about 19 chunks/s so far).

What I’m looking for

  1. A model with quality close to Qwen3-Embedding-4B, especially on
    paraphrased queries, that runs much faster on Apple Silicon.
  2. Or a small embedding model plus a reranker that gets there together.
  3. Any advice on the fastest runtime for these models on an M4
    (sentence-transformers on MPS, MLX, ONNX, llama.cpp).

Constraints: English only, runs locally in under about 3 GB of memory,
free to use.

Thanks for any suggestions!

Gopi Krishna Reddy Katkuri

nice eval, the paraphrased split is the one that matters and most people never build it.

two things i’d check before changing models. first, query prefixes: bge-small and bge-large expect "Represent this sentence for searching relevant passages: " in front of the query, and Qwen3-Embedding wants an instruct line before the query. leaving them out hurts paraphrased queries the most, so it might close part of that gap for free.

second, the 12 hours is a one-off. you embed the corpus once and then only new postings, while queries are a single short string each. if Qwen3-Embedding-0.6B lands near the 4B on your paraphrased set, run it overnight and stop there.

if you’d rather keep bge-small for indexing, retrieve the top 50 postings with it and rerank them with a small cross-encoder like bge-reranker-base. reranking 50 short pairs per query is quick on an M4 and tends to help most on exactly the paraphrased cases.

one caveat on the numbers: 112 questions is small, so a few points either way can be noise. look at which paraphrased questions each model gets wrong rather than only the percentages.