Skip to content

Update dependency sentence-transformers to v6 - #2170

Open
renovate[bot] wants to merge 1 commit into
devfrom
renovate/sentence-transformers-6.x
Open

renovate[bot] wants to merge 1 commit into
devfrom
renovate/sentence-transformers-6.x

Conversation

@renovate

@renovate renovate Bot commented Aug 26, 2026 •

Copy link
Copy Markdown
Contributor

ℹ️ Note

This PR body was truncated due to platform limits.

This PR contains the following updates:

Package Change Age Confidence
sentence-transformers ==5.7.0 → ==6.1.0 age confidence

Release Notes

huggingface/sentence-transformers (sentence-transformers)

v6.1.0: - Documentation and input handling improvements

Compare Source

This is a small release with refreshed inference benchmarks, clearer documentation, and a few improvements to multimodal input handling.

Install this version with

# Training + Inference
pip install sentence-transformers[train]==6.1.0

# Inference only, use one of:
pip install sentence-transformers==6.1.0
pip install sentence-transformers[onnx-gpu]==6.1.0
pip install sentence-transformers[onnx]==6.1.0
pip install sentence-transformers[openvino]==6.1.0

# Multimodal dependencies (optional):
pip install sentence-transformers[image]==6.1.0
pip install sentence-transformers[audio]==6.1.0
pip install sentence-transformers[video]==6.1.0

# Or combine as needed:
pip install sentence-transformers[train,onnx,image]==6.1.0

Changes

For existing multimodal dict inputs, message content now follows each input's key order, which can change embeddings or scores. Video-frame URLs without recognizable image extensions need an explicit frame wrapper to remain one video. See #​4037 for details.

All Changes

Full Changelog: huggingface/sentence-transformers@v6.0.1...v6.1.0

v6.0.1: - Restore the PyLate prefix on prompted checkpoints, 80 documented multi-vector models

Compare Source

This patch release fixes a multi-vector loading bug: PyLate checkpoints that carry both a [Q]/[D] prefix and a text prompt lost the prefix, so they were encoded without a marker they were trained with. It also grows the documented multi-vector model tables from 51 checkpoints to 80.

Install this version with

# Training + Inference
pip install sentence-transformers[train]==6.0.1

# Inference only, use one of:
pip install sentence-transformers==6.0.1
pip install sentence-transformers[onnx-gpu]==6.0.1
pip install sentence-transformers[onnx]==6.0.1
pip install sentence-transformers[openvino]==6.0.1

# Multimodal dependencies (optional):
pip install sentence-transformers[image]==6.0.1
pip install sentence-transformers[audio]==6.0.1
pip install sentence-transformers[video]==6.0.1

# Or combine as needed:
pip install sentence-transformers[train,onnx,image]==6.0.1

Compose the PyLate prefix with the saved prompt (#​3968)

PyLate supports a prefix as well as prompt text, and Sentence Transformers assumed that these were mutually exclusive. However, the following three checkpoints carry both a prefix and a prompt, and were trained with the prefix prepended to the prompt text:

PyLate trained these checkpoints on [CLS] [Q] search_query: ... and [CLS] [D] search_document: ..., but Sentence Transformers dropped the prefix and encoded them as [CLS] search_query: ... and [CLS] search_document: .... After the fix, the performance of the ColBERT-Zero model improved from 0.6569 NDCG@​10 to 0.6824 on the NanoBEIR benchmark. To my knowledge, only these 3 checkpoints used both a prefix and a prompt, so this is the only case where the bug would have affected you.

Forward task to routed modules in Router.preprocess (#​3967)

Router.forward has always passed task down to the routed module, but Router.preprocess did not, so anything a module does with the task at preprocessing time did nothing behind a Router: query_length and document_length caps, query_expansion, and the chat-template task keyword. No released checkpoint combines a Router with those settings, so this is a latent bug rather than one you are likely to have hit. It would have affected anyone building such a model themselves, with no error to indicate it.

Documentation

  • The multi-vector pretrained models tables grew from 29 text and 22 visual document retrieval checkpoints to 38 and 42 (#​3963, #​3969), with revision and trust_remote_code notes refreshed as upstream pull requests merged (#​3972).
  • Documented which Hub tag to filter on for each model type (#​3964), and linked the Multi-Vector Encoder blogposts from the docs (#​3966).
  • Added MultiVectorEncoder to the Agent Skill README and refreshed the SauerkrautLM scores (#​3955).
  • Corrected the minimum versions in the README (#​3952). It still recommended PyTorch 1.11.0+ and transformers v4.41.0+, where v6.0 requires PyTorch 2.2+ and transformers v5.0+. This is the text rendered as the PyPI project description.

All Changes

Full Changelog: huggingface/sentence-transformers@v6.0.0...v6.0.1

v6.0.0: - MultiVectorEncoder for ColBERT & late interaction models, transformers v5, float32 scoring, faster training & encoding

Compare Source

This major release introduces Multi-Vector Embedding models, also known as late interaction or ColBERT-style models, as a fourth model type alongside SentenceTransformer, CrossEncoder, and SparseEncoder. Going forward, you'll be able to use Sentence Transformers for training, inferencing, and interpreting Multi-Vector Embedding models.

It also modernizes the dependency floors to transformers v5, fixes a class of silent scoring bugs caused by half precision, and speeds up both training and encoding.

Install this version with

# Training + Inference
pip install sentence-transformers[train]==6.0.0

# Inference only, use one of:
pip install sentence-transformers==6.0.0
pip install sentence-transformers[onnx-gpu]==6.0.0
pip install sentence-transformers[onnx]==6.0.0
pip install sentence-transformers[openvino]==6.0.0

# Multimodal dependencies (optional):
pip install sentence-transformers[image]==6.0.0
pip install sentence-transformers[audio]==6.0.0
pip install sentence-transformers[video]==6.0.0

# Or combine as needed:
pip install sentence-transformers[train,onnx,image]==6.0.0

[!TIP]
Our Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers blogpost is an excellent place to learn about multi-vector models: loading the various checkpoint formats, encoding and scoring, plugging them into a search stack, running them on page images, and keeping the index affordable.

[!WARNING]
This is a major release with breaking changes. Upgrading from v5.x to v6.0 may require code updates. The changes marked 🚨 below are the ones most likely to affect you, and the Migration Guide has the full list. If you run into issues when upgrading, feel free to open an issue.

MultiVectorEncoder: ColBERT-style late interaction models (#​3794)

Sentence Transformers v6.0 introduces MultiVectorEncoder, for ColBERT-style late interaction retrieval. Where a regular embedding model compresses a whole text into one vector, a multi-vector model keeps one vector per token and scores query against document with the MaxSim operator. That preserves token-level matching information that a single vector has to average away, which usually means stronger retrieval at the cost of a bigger index. It is also the state of the art for visual document retrieval, where a text query is matched against page images directly, with no OCR step in between.

Any PyLate checkpoint and any Stanford-NLP ColBERT checkpoint loads straight into it, and colpali-engine models for visual document retrieval work too, through the same familiar API you already use for dense, sparse, and reranker models.

from sentence_transformers import MultiVectorEncoder

# Download from the 🤗 Hub
model = MultiVectorEncoder("lightonai/LateOn")

query_embeddings = model.encode_query(["Which planet is known as the Red Planet?"])
document_embeddings = model.encode_document([
    "Venus is often called Earth's twin because of its similar size and proximity.",
    "Mars, known for its reddish appearance, is often referred to as the Red Planet.",
    "Jupiter, the largest planet in our solar system, has a prominent red spot.",
    "Saturn, famous for its rings, is sometimes mistaken for the Red Planet.",
])

print(query_embeddings[0].shape)

# (12, 128)

scores = model.similarity(query_embeddings, document_embeddings)
print(scores)

# tensor([[10.7942, 11.1104, 10.9743, 11.0811]])

Mars wins, as it should, though notice how close the four scores are. That is normal for MaxSim: the scores often look similar, but the ranking is still exact. The blogpost explores this in more detail.

Note what you get back: a list of 2D tensors on the model device, one per input, each of shape (num_tokens, embedding_dim). Unlike dense embeddings, you cannot stack these into one rectangular tensor, because every input has its own token count. Pass convert_to_numpy=True for a list of numpy arrays instead, which is what you want once a corpus outgrows device memory.

Multi-vector models are also asymmetric: queries and documents go through different prefixes, different length caps, and different scoring masks. Unlike many dense models, where the two are interchangeable, encode_query and encode_document are required to get correct embeddings.

The MaxSim operator

Scoring uses MaxSim: for each query token, take its highest similarity against any document token, then sum those maxima across the query.

$$\text{MaxSim}(Q, D) = \sum_{Q_i \in Q} \max_{D_j \in D} Q_i \cdot D_j$$

You can read the operator as a soft alignment: every query token points at the one document token that best explains it, and the score is how well the document explains the query overall. The alignment does not have to be lexical, since the token embeddings are contextualized. But when an exact match does matter to you (a product code, a surname, a function name), MaxSim has a token sitting right there to match it, where a single-vector model had to fold it into an average.

Because MaxSim sums over query tokens, its magnitude scales with the query token count, so scores are not comparable across models with different query recipes. If you want scores on a bounded scale, use similarity_fn_name="meanmaxsim", which divides by the query token count and gives you an average cosine similarity in [-1, 1].

Scoring builds a 4-dimensional intermediate of every query token against every document token, which is the largest tensor in the operation. Every scoring function takes a chunk_elements budget that bounds it, defaulting to 100 million elements (roughly 400 MB in float32), so lower it if you run out of memory. Scores and gradients are bit-identical whatever you set it to. maxsim and maxsim_pairwise also take a device, which scores one chunk at a time on that device and moves each result straight back, letting you score a corpus larger than your VRAM on the GPU. Both are reachable through similarity, which forwards any extra keyword arguments to the scoring function:

scores = model.similarity(query_embeddings, document_embeddings, chunk_elements=1_000_000, device="cuda")

When training, pass the budget to the loss instead, with similarity_fct=partial(colbert_scores, chunk_elements=1_000_000). It chunks the document axis, so it composes with the loss-level score_mini_batch_size, which chunks the query axis.

Are they any good?

lightonai/LateOn and lightonai/DenseOn were trained by LightOn on the same data with the same ModernBERT backbone and the same 149M parameters, differing only in whether they keep one vector per token or pool down to one per document. Running both over all 13 NanoBEIR datasets isolates what that choice buys:

NanoBEIR dataset LateOn (multi-vector, 128d) DenseOn (dense, 768d)
MSMARCO 0.7194 0.6517
NQ 0.7810 0.7511
HotpotQA 0.9295 0.8802
FEVER 0.9702 0.9612
ClimateFEVER 0.4887 0.4846
DBPedia 0.6836 0.6748
QuoraRetrieval 0.9795 0.9687
Touche2020 0.5938 0.5673
ArguAna 0.5562 0.5660
NFCorpus 0.3949 0.3851
SciFact 0.7978 0.8057
SCIDOCS 0.4469 0.4484
FiQA2018 0.5871 0.6491
Mean 0.6868 0.6764

Late interaction wins on 9 of the 13 datasets and on the mean, by roughly one NDCG point. The four it loses (ArguAna, FiQA2018, SCIDOCS, and SciFact) are the shape of the tradeoff you should expect: a real gain in retrieval quality at the same model size, paid for in index footprint, rather than a universal win on every dataset. The same pair scores 57.22 against 56.20 on the full 15-dataset BEIR, a comparable gap, so the margin is not an artifact of the small benchmark.

That footprint is the real cost. One vector per token instead of one vector per document is a lot more vectors, only partly offset by the smaller dimension. Encoding 4,874 Natural Questions passages with lightonai/LateOn produced 608,414 token vectors, an average of 124.8 per passage:

Representation Vectors Dimensions float32 index
Dense, all-MiniLM-L6-v2 4,874 384 7.5 MB
Dense, gte-modernbert-base 4,874 768 15.0 MB
Multi-vector, LateOn 608,414 128 311.5 MB

That is about 42x the storage of the MiniLM index. Token Pooling cuts the vector count before any of that, real late interaction indexes compress heavily (the same vectors take 88 MB as a fast-plaid PLAID index), and using a multi-vector model as a reranker over a dense first stage avoids building an index at all.

Every checkpoint format loads

Multi-vector checkpoints have been published in several formats over the years. MultiVectorEncoder reads all of them, so loading looks the same whatever the model started life as:

from sentence_transformers import MultiVectorEncoder

# Native Sentence Transformers checkpoints. PyLate builds on the same schema,

# so any PyLate checkpoint loads identically
model = MultiVectorEncoder("lightonai/LateOn")
model = MultiVectorEncoder("mixedbread-ai/mxbai-edge-colbert-v0-17m")
model = MultiVectorEncoder("LiquidAI/LFM2-ColBERT-350M")

# Any Stanford-NLP ColBERT checkpoint, detected via the `HF_ColBERT` architecture

# marker. The inline projection weight and the recipe come from `artifact.metadata`
model = MultiVectorEncoder("colbert-ir/colbertv2.0")
model = MultiVectorEncoder("answerdotai/answerai-colbert-small-v1")

# transformers-native *ForRetrieval ports (ColPali, ColQwen2, ...)
model = MultiVectorEncoder("vidore/colqwen2-v1.0-hf")

# A bare transformer: a fresh random projection is appended, so training is required
model = MultiVectorEncoder("answerdotai/ModernBERT-base")

The recipe knobs that differ per checkpoint (marker prefixes for queries and documents, length caps, whether queries are padded out with [MASK] tokens, and which tokens are skipped when scoring documents) all live in the module configs, so print(model) shows you exactly what you loaded:

model = MultiVectorEncoder("colbert-ir/colbertv2.0")
print(model)
"""
MultiVectorEncoder(
  (0): Transformer({..., 'document_length': 180,
                    'query_expansion': {'strategy': 'fixed', 'attend': False, 'token': None, 'length': 32}})
  (1): Dense({'in_features': 768, 'out_features': 128, 'bias': False, ...})
  (2): MultiVectorMask({'skiplist_words': ['!', '"', '#', ...], 'skiplist_tasks': ['document'], ...})
  (3): Normalize({...})
)
"""

Following the design principle of the rest of the library, this behavior lives in swappable modules rather than in the model class: a Transformer producing contextualized token embeddings, a token-level Dense projecting each of them down, a MultiVectorMask deciding which tokens count during scoring, and a token-level Normalize.

Supported models

These are the checkpoints we test against directly, ranked by retrieval quality. The sentence-transformers tag on the Hub is the list that stays current, and for text retrieval in particular, any PyLate or Stanford-NLP ColBERT checkpoint loads whether or not it carries the tag yet. Where a revision is listed, pass it until the pull request on that repository is merged.

Text retrieval (29 models). NanoBEIR is the mean NDCG@​10 over the 13 NanoBEIR datasets, a fast proxy for English text retrieval quality. A - means the model was not evaluated on it, which is the case for the non-English models.

Model Parameters Dimensionality NanoBEIR Notes
lightonai/LateOn-regularized 149M 128 0.6897 -
lightonai/LateOn-hpool-regularized 149M 128 0.6876 -
lightonai/LateOn 149M 128 0.6868 -
LiquidAI/LFM2.5-ColBERT-350M 353M 128 0.6864 needs trust_remote_code=True
lightonai/mLateOn 307M 128 0.6851 -
VAGOsolutions/SauerkrautLM-Multi-ModernColBERT 149M 128 0.6741 -
lightonai/GTE-ModernColBERT-v1 149M 128 0.6720 -
topk-io/Iso-ModernColBERT 149M 128 0.6687 -
perplexity-ai/pplx-embed-v1-late-0.6b 596M 128 0.6662 needs trust_remote_code=True
VAGOsolutions/SauerkrautLM-Multi-Reason-ModernColBERT 149M 128 0.6616 -
lightonai/ColBERT-Zero 149M 128 0.6569 -
answerdotai/answerai-colbert-small-v1 33M 96 0.6550 -
mixedbread-ai/mxbai-edge-colbert-v0-32m 32M 64 0.6524 -
LiquidAI/LFM2-ColBERT-350M 353M 128 0.6441 -
mixedbread-ai/mxbai-edge-colbert-v0-17m 17M 48 0.6407 -
lightonai/colbertv2.0 110M 128 0.6201 -
lightonai/LateOn-Code 149M 128 0.6169 -
lightonai/Agent-ModernColBERT 149M 128 0.6164 -
lightonai/Reason-ModernColBERT 149M 128 0.6078 -
colbert-ir/colbertv2.0 110M 128 0.6053 -
VAGOsolutions/SauerkrautLM-Reason-EuroColBERT 212M 128 0.6039 -
VAGOsolutions/SauerkrautLM-EuroColBERT 212M 128 0.5965 -
antoinelouis/colbert-xm 853M 128 0.5915 -
mixedbread-ai/mxbai-colbert-large-v1 335M 128 0.5733 -
lightonai/LateOn-Code-edge 17M 48 0.5274 -
NeuML/biomedbert-base-colbert 110M 128 0.4320 -
yjoonjang/colbert-ko-v1 149M 128 - -
ytu-ce-cosmos/turkish-colbert 111M 256 - -
samheym/GerColBERT 110M 128 - -

Visual document retrieval (22 models). These embed page images as documents and text as queries. NanoViDoRe is the equivalent proxy over the ViDoRe benchmark subsamples.

Model Parameters NanoViDoRe Notes
webAI-Official/webAI-ColVec1.1-8b 8.4B 0.6580 needs trust_remote_code=True
webAI-Official/webAI-ColVec1.1-4b 4.5B 0.6520 needs trust_remote_code=True
tencent/EVIE-Preview-4.5B 4.54B 0.6405 -
TomoroAI/tomoro-colqwen3-embed-8b 8.8B 0.6206 needs trust_remote_code=True
TomoroAI/tomoro-colqwen3-embed-4b 4.4B 0.6019 needs trust_remote_code=True
vidore/colqwen2.5-v0.2 3.8B 0.5402 -
vidore/colqwen2.5-v0.1 3.8B 0.5395 -
vidore/colqwen-omni-v0.1 4.4B 0.5309 -
vidore/colpali-v1.3 2.9B 0.4802 -
vidore/colpali-v1.3-hf 2.9B 0.4793 -
vidore/colpali-v1.2 2.9B 0.4691 -
vidore/colqwen2-v1.0 2.2B 0.4685 -
vidore/colqwen2-v0.1 2.2B 0.4526 -
vidore/colpali 2.9B 0.4516 -
vidore/colpali-v1.1 2.9B 0.4314 -
vidore/colsmolvlm-v0.1 2.1B 0.4054 -
vidore/colpali-hard-v1.1 2.9B 0.3949 -
vidore/colSmol-500M 507M 0.3459 -
vidore/colSmol-256M 256M 0.2673 -
ModernVBERT/colmodernvbert 252M 0.2632 -
vidore/colpali-v1.2-hf 2.9B - -
vidore/colqwen2-v1.0-hf 2.2B - -

Note that NanoBEIR and NanoViDoRe are small benchmarks, so their scores are not a substitute for evaluating on your own data, which is always the right way to pick a model.

Visual, audio, and video document retrieval

Late interaction is the state of the art for visual document retrieval: matching a text query against page images, with charts, tables, and layout intact, and no OCR step. This is what the ColPali family of models does, and those checkpoints run through the same API. Image documents are passed as URLs, local paths, or PIL images:

from sentence_transformers import MultiVectorEncoder

model = MultiVectorEncoder("vidore/colqwen2.5-v0.2")

queries = [
    "What is the variable represented on the y-axis of the graph?",
    "Total outlay is maximum in which year?",
]
images = [
    "https://huggingface.co/datasets/sentence-transformers/example-documents/resolve/main/doc1.jpg",
    "https://huggingface.co/datasets/sentence-transformers/example-documents/resolve/main/doc2.jpg",
    "https://huggingface.co/datasets/sentence-transformers/example-documents/resolve/main/doc3.jpg",
    "https://huggingface.co/datasets/sentence-transformers/example-documents/resolve/main/doc4.jpg",
]

# A page yields far more vectors than a query: one per image patch
query_embeddings = model.encode_query(queries)
document_embeddings = model.encode_document(images)
print(query_embeddings[0].shape, document_embeddings[0].shape)

# (25, 128) (755, 128)

scores = model.similarity(query_embeddings, document_embeddings)
print(scores)

# tensor([[13.8672, 12.3115, 12.1670, 11.0293],
#         [ 7.2012, 14.7207,  6.9414,  6.9746]])

The code is unchanged from the text case. Underneath, the processor handles the visual prompt and the image patches, and MaxSim scores query text tokens against document image patches. Page images are not the only non-text modality either: text, images, audio, and video are all accepted, and a checkpoint supports whichever of those its processor does, which model.modalities reports.

Because MaxSim is a sum of per-query-token maxima, a ranking decomposes exactly: every point of a document's score belongs to one query token and one document token. The new sentence_transformers.multi_vector_encoder.interpretability module overlays that decomposition onto the page as the standard ColPali heatmap, either aggregated over the query or one map per query token.

Token pooling

If the index footprint worries you, the most effective knob is to store fewer token vectors. HierarchicalTokenPooling implements the token pooling technique from Clavié, Chaffin, and Adams: it clusters each document's token vectors with Ward linkage on cosine similarity and replaces each cluster with its mean, keeping roughly 1 / pool_factor of the tokens.

from sentence_transformers import MultiVectorEncoder
from sentence_transformers.multi_vector_encoder.modules import HierarchicalTokenPooling

model = MultiVectorEncoder("lightonai/LateOn")
pooling = HierarchicalTokenPooling(pool_factor=2)

# 1. Per encode call
document_embeddings = model.encode_document(documents, token_pooling=pooling)

# 2. Standalone, on embeddings you already have saved
pooled = pooling.pool(document_embeddings)

# 3. Baked into the model, so every consumer of the checkpoint gets pooled documents
model.append(HierarchicalTokenPooling(pool_factor=2))
model.save_pretrained("my-pooled-colbert")

By default, pooling applies to documents only, since queries are short and are the side you cannot afford to distort. On the Natural Questions corpus above, the reduction tracks pool_factor closely:

pool_factor Token vectors Reduction float32 index
1 (off) 608,414 1.00x 311.5 MB
2 305,438 1.99x 156.4 MB
3 204,407 2.98x 104.7 MB
4 153,936 3.95x 78.8 MB

The original experiments measured the retrieval cost of this on BEIR and found very little of it: 100.6% of the unpooled performance on average at pool_factor=2, and 99.0% at pool_factor=3. How much it costs on your data is corpus-specific, so measure it with an evaluator before you settle on a factor.

Update Stats

Introducing MultiVectorEncoder has been one of the largest updates to Sentence Transformers, introducing all of the following:

Resources

🚨 transformers v5, torch 2.2, and new dependency floors (#​3794)

Sentence Transformers v6.0 requires transformers v5. The v4.x compatibility branches have been removed, which is what allows the new modality handling, chat template support, and unpadding paths to be relied upon rather than feature-detected. The floors that moved:

Dependency v5.7.0 v6.0.0
transformers >=4.41.0,<6.0.0 >=5.0.0,<6.0.0
huggingface-hub >=0.23.0 >=1.3.0,<2.0.0
torch >=1.11.0 >=2.2
numpy >=1.20.0 >=1.24.0
scikit-learn >=0.22.0 >=1.1.0
typing_extensions >=4.5.0 >=4.10.0
datasets (train) >=2.0.0 >=2.16.0
accelerate (train) >=0.20.3 >=1.3.0
optimum-intel[openvino] unpinned >=2.0.0

requires-python is unchanged at >=3.10. Note that multi-GPU training with streaming (IterableDataset) datasets needs accelerate>=1.13.0 in practice.

🚨 Higher-precision scoring (#​3892, #​3893, #​3924, #​3926)

Half precision ties too many scores together to rank with. Three separate places where that mattered are now computed in float32.

Reranker scores are the big one. CrossEncoder.predict (and rank) now upcast the logits to float32 before applying the activation function. A sigmoid in bfloat16 saturates and collapses the top candidates onto a handful of tied values, which randomizes their order. Measured on cross-encoder/ettin-reranker-32m-v1 in bfloat16 over three NanoBEIR datasets with 100 candidates per query:

Metric v6.0.0 v5.7.0
NanoBEIR mean NDCG@​10 0.6795 0.1849
NanoBEIR mean MRR@​10 0.6797 0.3986
Unique scores over 15,040 pairs 710 270

NanoMSMARCO NDCG@​10 alone goes from 0.0965 to 0.7093. If you run a half precision reranker with the default sigmoid activation, its ranking was essentially randomized before this release. Models using activation_fn=nn.Identity() (raw logits) were unaffected, as bf16 logits keep enough relative spacing.

Similarity scores from model.similarity / similarity_pairwise and the cos_sim family are now computed in float32 for float16 and bfloat16 embeddings. With 10,000 realistic cosine scores (mean 0.7, standard deviation 0.05), float32 keeps 9,983 distinct values where float16 keeps 593 and bfloat16 keeps just 93. bfloat16 can represent only 129 distinct values in the whole of [0.5, 1.0).

MaxSim sums over query tokens, reaching magnitudes where the bfloat16 grid is 0.125 wide, so maxsim and maxsim_pairwise accumulate the per-token maxima in float32 and always return float32 scores. The 4-dimensional scoring intermediate stays in the input dtype, so this does not change peak memory.

Note that encode() output dtypes are unchanged. Only the scoring step is upcast. For CrossEncoder.predict, the returned dtype changes only with convert_to_tensor=True or convert_to_numpy=False, as the default numpy output was already float32.

Separately, the multi-vector bf16 benchmarks were re-measured under this float32 accumulation (#​3924). Most of the previously reported bf16 quality drop came from the scoring accumulation rather than from the embeddings: plain bf16 now sits at 99.0% of fp32 retrieval quality (was 95.0%), and bf16 with FlashAttention-2 is indistinguishable from fp32 at 99.96% (was 97.9%).

🚨 Other breaking changes (#​3794, #​3927, #​3935)

  • similarity and similarity_pairwise are methods, not properties. Calls like model.similarity(embeddings1, embeddings2) work unchanged, but assigning a custom function to model.similarity is no longer supported: it now silently shadows the method where it previously raised an AttributeError. Set model.similarity_fn_name = "dot" instead, which updates both. Note also that model.similarity.__name__ is now "similarity" rather than the resolved function name, which affected loss get_config_dict() output and generated model cards. The new sentence_transformers.util.similarity_fct_name() resolves it properly and the losses use it.
  • A bare list of chat message dictionaries is now one conversation. model.encode([{"role": "user", ...}, {"role": "assistant", ...}]) produces one embedding, where v5.x read it as a batch of two inputs. Wrap each conversation in its own list to encode a batch: model.encode([[msg1], [msg2]]). This applies to SentenceTransformer, SparseEncoder, and MultiVectorEncoder. CrossEncoder is unaffected.
  • Custom module classes require trust_remote_code=True (#​3935). Loading a model whose modules.json references a class outside sentence_transformers executes third-party code, and a local directory no longer implies trust. This closes the bypass reported in #​3801 and completes the deprecation cycle announced in v5.6 and v5.7. Unmet, it raises a ValueError naming the class and pointing at the repository or local path to inspect. Trainer checkpoint reloading (load_best_model_at_end, resume_from_checkpoint) keeps working for programmatically built models without the flag.
  • quantize_embeddings returns a list of per-input matrices when given a list of 2D arrays, where it previously stacked them into one 3D array. Update callers that indexed the stacked array. An empty list now returns [] instead of raising, and a (0, dim) matrix returns a correctly shaped empty result.
  • Multi-process encode(pool=..., precision="int8") now quantizes once after merging the worker results, so the calibration ranges match single-process encoding. Quantized indexes built with v5.x multi-process encoding are not bit-compatible and should be regenerated. Peak memory is higher, because the full float32 matrix is materialized before quantization.
  • CrossEncoder.rank returns Python floats (#​3927) as its "score" values, where it previously returned numpy.float32 scalars or 0-dimensional tensors. The results are directly JSON serializable, matching semantic_search. convert_to_numpy and convert_to_tensor on rank are now deprecated no-ops: call predict directly if you want an array or a tensor. Beyond the cleaner output, this avoids a device synchronization per comparison when sorting, which took 212ms for 1000 CUDA scalars against 0.089ms for Python floats.
  • Normalize moved to sentence_transformers.base.modules. Existing models load fine and silently, but a model saved by v6.0 with a Normalize module cannot be loaded by Sentence Transformers older than v6.0.
  • Cross-family conversion no longer inherits inference settings. Loading a SentenceTransformer checkpoint as a CrossEncoder (or any other such conversion) no longer picks up the source's prompts, default_prompt_name, similarity_fn_name, truncate_dim, or activation_fn, as those describe a model you are not loading. A reranker's default prompt being prepended to every encode call was the motivating case. Explicit keyword arguments still win. These conversions are now also logged at warning level, so they are visible at default verbosity.
  • SimilarityFunction.possible_values() now includes "maxsim" and "meanmaxsim". Setting an unsupported similarity_fn_name on SentenceTransformer or SparseEncoder raises immediately rather than failing later, and a new SUPPORTED_SIMILARITY_FN_NAMES class attribute documents what each model type accepts.
  • Automatically generated model cards now open with "It maps inputs to a N-dimensional dense vector space" rather than "It maps sentences & paragraphs to ...", since the model could be handling other modalities.

Faster training and encoding (#​3938, #​3794)

Multi-column losses now run one forward pass over merged columns (#​3938). A training batch arrives as one feature dict per column (anchor, positive, negative_1, and so on), and the classic pattern runs the model once per column. The SentenceTransformer and SparseEncoder losses now pad and concatenate the like-width candidate columns into a single batch, keeping the anchor on its own forward pass since a 12-token query padded into 256-token documents costs more than it saves:

Configuration v5.7.0 v6.0.0
Natural Questions with 5 hard negatives 316.2s 250.7s (1.26x)
AllNLI triplets 53.8s 44.2s (1.22x)

Loss trajectories match, up to dropout sampling. Losses fall back to per-column forward passes whenever the columns cannot be merged safely, for example with differing feature keys, disagreeing prompts or router tasks, or flattened Flash Attention inputs. The cached losses keep using GradCache, and AdaptiveLayerLoss opts out.

Backend benchmarks were re-measured for all four model types, with new Flash Attention columns and rewritten recommendations. For SentenceTransformer, float16 with Flash Attention and unpadding is now the fastest GPU configuration at 3.87x over float32, and ONNX on GPU is no longer recommended for short texts as float16 now beats it. For CrossEncoder, Flash Attention is explicitly not recommended, as unpadding does not apply to classification heads. For SparseEncoder, plain float16 remains the recommendation even though FA2 unpadding is now supported. See Speeding up Inference for the flowcharts.

Models can declare their dependency versions (#​3934)

Model authors can now record which package versions their checkpoint needs, and loading verifies them up front instead of failing in a confusing way later. Add a requirements mapping to config_sentence_transformers.json, using PEP 440 specifiers:

{
    "model_type": "SentenceTransformer",
    "requirements": {
        "transformers": ">=5.15",
        "peft": {
            "specifier": ">=0.18,<0.20",
            "reason": "Older versions ignore the key_mapping, which silently randomizes the adapter weights."
        }
    }
}

Loading that model in an environment that does not satisfy it raises an ImportError listing every unmet requirement at once, with the optional reason included and a ready-to-run install command:

The model 'tomaarsen/my-model' requires:
- transformers>=5.15, but transformers==5.4.0 is installed.
- peft>=0.18,<0.20, but peft==0.17.0 is installed. Older versions ignore the key_mapping, which silently randomizes the adapter weights.
Install compatible versions with:
    pip install -U

> ❗ **Important**
> 
> ✂ PR body was truncated to here.


</details>

---

### Configuration

📅 **Schedule**: (in timezone America/New_York)

- Branch creation
  - At any time (no schedule defined)
- Automerge
  - At any time (no schedule defined)

🚦 **Automerge**: Disabled by config. Please merge this manually once you are satisfied.

♻ **Rebasing**: Whenever PR becomes conflicted, or you tick the rebase/retry checkbox.

🔕 **Ignore**: Close this PR and you won't be reminded about this update again.

---

 - [ ] <!-- rebase-check -->If you want to rebase/retry this PR, check this box

---

This PR was generated by [Mend Renovate](https://mend.io/renovate/). View the [repository job log](https://developer.mend.io/github/arthur-ai/arthur-engine).
<!--renovate-debug:eyJjcmVhdGVkSW5WZXIiOiI0NC40Ni4wIiwidXBkYXRlZEluVmVyIjoiNDQuMTAzLjAiLCJ0YXJnZXRCcmFuY2giOiJkZXYiLCJsYWJlbHMiOlsibWFqb3IiXX0=-->

@renovate renovate Bot added the major label Aug 26, 2026
@renovate
renovate Bot requested review from notthattal and ntatsumi August 26, 2026 12:21
@github-actions

github-actions Bot commented Aug 26, 2026 •

Copy link
Copy Markdown
Contributor

Coverage

Warning

Your comment is too long (maximum is 65536 characters), so the coverage report was not added. See the job log for how to reduce it.

Tests Skipped Failures Errors Time
2090 7 💤 0 ❌ 0 🔥 19m 0s ⏱️

@renovate
renovate Bot force-pushed the renovate/sentence-transformers-6.x branch 7 times, most recently from 1f3af67 to d271de6 Compare August 27, 2026 05:27
@ntatsumi
ntatsumi enabled auto-merge August 27, 2026 05:37
@renovate
renovate Bot force-pushed the renovate/sentence-transformers-6.x branch 2 times, most recently from 17aa0bd to a9f10af Compare August 27, 2026 06:57
@ntatsumi
ntatsumi removed their request for review August 27, 2026 13:36
@renovate
renovate Bot force-pushed the renovate/sentence-transformers-6.x branch 15 times, most recently from 617f2a6 to 7ebe8c7 Compare August 30, 2026 13:03
@renovate
renovate Bot force-pushed the renovate/sentence-transformers-6.x branch 27 times, most recently from 003fbb2 to c33d363 Compare September 9, 2026 06:34
@claude

claude Bot commented Sep 11, 2026

Copy link
Copy Markdown
Contributor

🤖 Autofix self-review

Verdict: fix is sound, all invariants hold. The sentence-transformers 6.0.1 update lands untouched; the red check was a pre-existing wall-clock flake.

Root cause

The failing check was run-genai-engine-unit-tests:

FAILED tests/unit/routes/tasks/test_agentic_prompt_routes.py::test_get_all_prompts_includes_deleted_prompts
  - assert 1.148429 < 1

test_get_all_prompts_includes_deleted_prompts saves two versions of a prompt with two
sequential POSTs, then asserted:

assert abs((created - latest).total_seconds()) < 1

created_at is MIN(created_at) and latest_version_created_at is MAX(created_at) over the
prompt's versions (base_llm_repository.py:562-568).
With two versions those are exactly the timestamps of the test's own two POSTs, so the assertion
did not measure any property of the API — it measured how fast the runner served two HTTP
requests
. The audit log in the failed run shows the two saves landing at 07:42:10.625 and
07:42:11.774: 1.15s apart, so the 1-second budget was blown.

This is not the dependency update

  • The same dependency change passed this exact job three times on this branch:
    34570621770 (21m41s), 34567579800 (21m26s),
    34490820503 (20m15s). Only run 34574503029 (23m39s) went red.
  • The prompt-save path never touches sentence_transformers; src/ is byte-identical to the
    branch HEAD in this fix.
  • The same runner slowness hit unrelated branches the same night — the dev merge queue run
    34539985363 failed on a different timing assertion
    (TimeoutError: Experiment ... did not complete within 30 seconds).

Run duration varies 20–24 min across runs of identical code, i.e. runner noise, not a regression
from the bump.

Fix

Both tests now pin the two aggregate timestamps to the version rows they are derived from,
instead of to how fast the saves ran:

assert created == version_created_at[1]
assert latest == version_created_at[2]

Two files, one hunk each — test_agentic_prompt_routes.py (the test that failed) and
test_llm_eval_routes.py (byte-identical race, same 1s budget, same suite). I fixed the sibling
deliberately rather than opportunistically: it is a coin-flip away from turning this PR red again
on the retry. The other two sites of this pattern (test_agentic_prompt_routes.py:137,
test_llm_eval_routes.py:803) are not touched — they cover single-version assets where
MIN == MAX, so the difference is always exactly 0.0 and they cannot flake.

Invariant 2: the check got stronger, not weaker

Net +3 / −1 assertions, and the replacement is strictly more discriminating. Mutation-tested:
flipping sa.func.min → sa.func.max for created_at in base_llm_repository.py makes both
tests fail —

assert datetime(2026, 9, 11, 8, 1, 8, 889956) == datetime(2026, 9, 11, 8, 1, 8, 879404)

— a 10µs error that the old < 1 tolerance passed happily. The new assertions catch a real
contract break the old one could not see. No skip/xfail, no # type: ignore, no deleted
assertions, no coverage reduction (88%, floor 79%).

Verification

Step Result
Reproduce root cause (1.2s delay injected between the two POSTs) ❌ assert 1.212019 < 1 — same shape as CI's 1.148429 < 1
Both tests with the fix, delay still injected ✅ 2 passed
Mutation test (min→max) against fixed assertions ❌ both fail as intended
Full suite pytest -m unit_tests ✅ 1993 passed, 7 skipped, exit 0 (18m25s), coverage 88%
isort --check / autoflake --check / black --check / mypy on src ✅ all four clean
black/isort/autoflake on the two edited test files ✅ clean
.github/scripts/check-lockfile-drift.sh ✅ all 4 uv projects agree with their manifests
Installed version in the synced venv ✅ sentence_transformers.__version__ == 6.0.1

Unit tests run against SQLite via a dependency override, and the CI job declares no postgres
service — so the local run exercises the same backend as CI.

Invariants

  1. Update lands — ✅ no dependency constraint touched. pyproject.toml keeps
    sentence-transformers==6.0.1; venv resolves to 6.0.1. Nothing capped, downgraded or excluded.
  2. No check weakened — ✅ strengthened, mutation-proven above.
  3. No behaviour change — ✅ src/ untouched; git status lists only the two test files.
  4. Lockfiles generated, not authored — ✅ no lockfile edited; drift check green.
  5. Release-age soak respected — ✅ minimumReleaseAge / exclude-newer untouched.
  6. CI rules off limits — ✅ nothing under .github/ or renovate.json modified.
  7. Constraint edits declared — ✅ none to declare. No dependency constraint changed, so
    no autofix: dependency-realignment label applies.

Caveat for reviewers

This makes two flaky assertions deterministic; it does not make the suite immune to runner
slowness in general. The 30s experiment-completion timeout that failed on dev overnight is the
same class of problem and is still there — out of scope here, but worth a separate look.

🤖 Generated with Claude Code

@coderabbitai

coderabbitai Bot commented Sep 21, 2026 •

Copy link
Copy Markdown

Review Change StackReview Change Stack

Understand this PR’s impact

Explore downstream dependencies and potential security impact with Blast Radius.

View blast radius →

Important

Review skipped

Review was skipped as selected files did not have any reviewable changes.

⛔ Files ignored due to path filters (1)
  • genai-engine/uv.lock is excluded by !**/*.lock
⚙️ Run configuration
  • Configuration used: defaults
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: c7dede89-6728-4e71-b148-249bc9da26e6
📥 Commits

Reviewing files that changed from the base of the PR and between 826cf6d and d26c683.

⛔ Files ignored due to path filters (1)
  • genai-engine/uv.lock is excluded by !**/*.lock

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: f8114903-2c1c-4ea6-be9d-d335edfc34c8

📥 Commits

Reviewing files that changed from the base of the PR and between 2f2d846 and f32dc79.

⛔ Files ignored due to path filters (1)
  • genai-engine/uv.lock is excluded by !**/*.lock
📒 Files selected for processing (1)
  • genai-engine/pyproject.toml

Included review availability: Your plan provides up to 10 included reviews per hour; 4 remain after this review.


📝 Walkthrough

Walkthrough

The pull request updates the pinned sentence-transformers dependency from 5.7.0 to 6.1.0.

Changes

Dependency Update

Layer / File(s) Summary
Update sentence-transformers pin
genai-engine/pyproject.toml
The dependency pin changes from 5.7.0 to 6.1.0.

Priority: ⬇️ Low

Estimated code review effort: 1 (Trivial) | ~2 minutes

Change: Other

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the main change: updating the sentence-transformers dependency to version 6.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check. Docstring coverage is scoped to functions touched by this diff. Analyzed 0 functions across 0…
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR

Comment @coderabbitai help to get the list of available commands.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Development

Successfully merging this pull request may close these issues.

0 participants