Skip to content

fix: load official SenseVoice ONNX models - #102

Open
LauraGPT wants to merge 1 commit into
lanbinleo:mainfrom
LauraGPT:codex/fix-sensevoice-onnx-bootstrap-20260812
Open

LauraGPT wants to merge 1 commit into
lanbinleo:mainfrom
LauraGPT:codex/fix-sensevoice-onnx-bootstrap-20260812

Conversation

@LauraGPT

@LauraGPT LauraGPT commented Aug 12, 2026 •

Copy link
Copy Markdown

Fixes #101.

Summary

  • detect model_quant.onnx and pass quantize=True to funasr_onnx.SenseVoiceSmall
  • keep custom non-quantized model.onnx directories working with quantize=False
  • validate the required config, CMVN, tokenizer, and ONNX files before importing/loading the runtime
  • document the exact two-repository ModelScope preparation flow in Chinese and English

The official iic/SenseVoiceSmall-onnx repository ships model_quant.onnx but does not include chn_jpn_yue_eng_ko_spectok.bpe.model, while the existing loader used the constructor default quantize=False. This made the documented SenseVoice path fail even after downloading the official ONNX repository.

Validation

  • 59 passed with the locked project environment (uv sync --extra web)
  • focused quantized, non-quantized, and missing-tokenizer regression tests
  • Ruff check and format check on changed Python files
  • Python compile check and git diff --check
  • verified the actual funasr-onnx==0.4.1 wheel constructor supports quantize: bool = False
  • verified both referenced ModelScope repositories are reachable

Validation Refresh (2026-09-29)

Rechecked unchanged head 27006d703898c7447369b745fb052f66b5618c44 against main 0248c84593bf66391dc2e38b4583ff72e3eae456 in the project's locked Linux/Python 3.12.3 environment:

  • uv sync --frozen --extra web --extra sensevoice completed successfully, including the actual optional runtime dependencies. No lockfile changes.
  • Actual funasr_onnx.SenseVoiceSmall and postprocessor imports succeed with funasr-onnx==0.4.1, torch==2.11.0, and onnxruntime==1.24.4.
  • uv pip check --python .venv/bin/python: all 95 installed packages compatible.
  • uv run --frozen --no-sync --extra web --extra sensevoice pytest -q -ra: 59 passed, 6 dependency warnings, no skipped tests.
  • git diff --check: passed; repository source and PR head unchanged.
  • Both the minimum supported 0.4.0 wheel and locked 0.4.1 wheel support the quantize constructor argument. Downloaded library wheel digests were checked against PyPI metadata.
  • Fresh official ModelScope file listings confirm the documented split: the ONNX repository contains model_quant.onnx and supporting config/CMVN/token files, while the tokenizer .bpe.model must come from iic/SenseVoiceSmall. Config and CMVN metadata hashes match across the two repositories.

These are environment, import, unit-test, runtime-contract, and remote metadata checks. No model weights or tokenizer artifacts were downloaded, no model was loaded, and no audio inference or GPU validation was performed. Earlier Ruff/format/compile checks above are historical, not rerun as part of this refresh.

Separate Pre-existing ITN Gap

Issue #105 tracks a separate issue already present on main: the adapter sends use_itn, but locked funasr-onnx==0.4.1 reads textnorm instead. A no-weights probe against the actual runtime selects woitn / 15 for both adapter settings, while an explicit textnorm="withitn" control selects 14. Existing passing tests do not cover that transcription argument contract. This PR does not fix ITN selection; its scope remains model-directory preparation and quantization.

Signed-off-by: LauraGPT <LauraGPT@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

SenseVoice 官方模型仓库无法开箱即用:quantize 参数 + 模型文件缺失

1 participant