accounts/docs/bench_embedding.md
Haitao Pan 84cb83933d Add benchmarking scripts and configs:
- bench_embedding.sh / bench_ollama.sh for Ollama & embedding API tests
- hf_embedding_bench.py for HF model performance
- models.txt / models-emb.txt for test configs
- docs in bench_embedding.md / bench_ollama.md
2025-08-13 13:12:34 +08:00

34 lines
1.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

建一个模型清单:
echo bge-m3:latest > models-emb.txt
跑评估(限定 1024 维,更贴合 pgvector(VECTOR(1024))
bash docs/bench_embedding.sh --input_config models-emb.txt --require-dim 1024
# 如需导出 CSV
CSV_OUTPUT=1 bash docs/bench_embedding.sh --input_config models-emb.txt --require-dim 1024
常用命令(以 1024 维 BAAI/bge-m3 为例)
# MPSMac+ FP16128 token 文本,批量 32共 2048 样本
python hf_embedding_bench.py --model BAAI/bge-m3 --device mps --dtype float16 \
--batch-size 32 --seq-len 128 --num-samples 2048 --warmup 10
# CUDA + BF16若支持批量 64
python hf_embedding_bench.py --model BAAI/bge-m3 --device cuda --dtype bfloat16 \
--batch-size 64 --seq-len 128 --num-samples 4096 --warmup 20 --compile
# CPU 基线
python hf_embedding_bench.py --model BAAI/bge-m3 --device cpu --dtype float32 \
--batch-size 16 --seq-len 128 --num-samples 512
小提示(让数字更准)
CUDA/MPS 下会自动 synchronize()计时更准确forward 使用 autocastfp16/bf16提升吞吐。
--seq-len 是近似长度,脚本会用 tokenizer 截断到 --max-length默认 512
tokens/s 是按 attention_mask 实际非 padding 的 token 数统计的。
你要严格对齐 1024 维:看脚本输出的 Output dim: 是否为 1024bge-m3 会是 1024和你 PG 表 VECTOR(1024) 一一对应。