Codec 2 1.2.0 · 可审计模型档案¶
当前结论:D0–D5 passed;下一步是 Active D0-D5 research moves to DAC. Codec 2 D6/D7 run only if it is selected as a Pareto anchor, teacher, deployment target or original-baseline comparator.。 每个 Gate 只建立该 Gate 明示的证据;D5 passed 表示冻结共同 benchmark 已完整归档和审计,不表示存在全局赢家或 D6/D7 已完成。 D0 只证明身份、来源和本地 artifact 一致;后续 Gate 的通过不会反向扩大 D0 的证据边界。
Gate 状态¶
| Gate | 状态 |
|---|---|
| D0 | passed |
| D1 | passed |
| D2 | passed |
| D3 | passed |
| D4 | passed |
| D5 | passed |
| D6 | not_started |
| D7 | not_started |
D0 · 身份、标准与本地 artifact¶
| 项目 | 冻结身份 |
|---|---|
| 主论文/规范/资料 | Codec 2 Algorithm Description · CODEC2-MANUAL@310777b1 |
| 规范更新 | — |
| 官方实现 | https://github.com/drowe67/codec2 · 1.2.0 / 06d4c11e699b |
| 本地环境 | Darwin 26.5 · arm64 · Apple M5 |
| 学习参数/模型状态 | 10752 / 43008 bytes(传统 codec,无 checkpoint) |
| 已冻结运行 artifact | 5 个,共 1,664,608 bytes |
| 当前公共点 | 1.2 kbps, 2.4 kbps, 3.2 kbps · application=narrowband speech · signal hint=not applicable · fixed mode bit allocation · mode-dependent 40 (1200) or 20 (2400/3200) ms |
冻结范围与明确排除¶
-
不纳入当前身份:
- FreeDV modems, FEC, framing and radio-channel behavior
- 700C, 1300, 1400 and 1600 modes not selected for the common benchmark
- optional LPCNet/FreeDV 2020 builds and external ML experiments
- post-1.2.0 RADE or current-main algorithms
- the living 2025 algorithm manual as an exact byte-level description of every 2023 release path
Primary-source manifest¶
| ID | 作用 | bytes | SHA-256 |
|---|---|---|---|
| CODEC2_AUTHOR_PAGE | author-stated purpose, range, application, informal demo/evaluation boundary and source links | 24,778 | e8765e9f9847c8fcaedebcaccffbd3ba8b7aa69275ffa4d6253b5719c95ade9e |
| CODEC2_RELEASE_README | selected release build, CLI, raw 8-kHz input, test and repository-boundary description | 9,268 | da7db4e687ba20ce687e8e7187f59bb29f95c40c303806b710fd97027d924909 |
| CODEC2_RELEASE_API | release-exact mode constants and encode/decode/frame-query interface | 4,104 | b8d2c2a18f03aff240f4d5799ba4010ac817509304a225647dd40ef4ad9c4857 |
| CODEC2_ALGORITHM_MANUAL_CURRENT | author-maintained conceptual source for D1, explicitly time-conditioned and not assumed release-exact | 316,470 | 957f9b8163c4bd34b573593c035bcee79204450a373f027377c36a48e1050f49 |
Source releases and licenses¶
| Artifact | revision | bytes | SHA-256 | license |
|---|---|---|---|---|
| codec2-1.2.0.tar.gz | 1.2.0 / 06d4c11e699b0351765f10398abb4f663a984f36 |
7,688,549 | cbccae52b2c2ecc5d2757e407da567eb681241ff8dadce39d779a7219dbcf449 |
LGPL-2.1-only |
| codec2--1.2.0.arm64_tahoe.bottle.tar.gz | Homebrew codec2 1.2.0 arm64_tahoe bottle rebuild 0 |
736,664 | efb5e3717cb204bc90d294e0353ed8e610795cd64950a17755eb74add01fcc95 |
LGPL-2.1-only |
Installed encoder/decoder identity¶
| Artifact | bytes | SHA-256 | resolved path |
|---|---|---|---|
c2enc |
69,328 | 160e311db04a7ff18531fa3458fa321ba114a8e62dd23ddee321ffe43afc524b |
/opt/homebrew/Cellar/codec2/1.2.0/bin/c2enc |
c2dec |
70,176 | 28ff84b495943cda65513f1c58c19b82e84e2aa2873d7aa2bd675c00433a8975 |
/opt/homebrew/Cellar/codec2/1.2.0/bin/c2dec |
c2demo |
68,720 | 9169294561ff499e34171d5c36fa27f3d14ee094f3d5a22c9329929d10dd8911 |
/opt/homebrew/Cellar/codec2/1.2.0/bin/c2demo |
c2sim |
89,376 | 5b733c2479f7dff5f6aa2a6273eca7f2113db764020b62ca5646943a5740316e |
/opt/homebrew/Cellar/codec2/1.2.0/bin/c2sim |
libcodec2 |
1,367,008 | d129ce4c9c3cfedb276087d0d612da56162bbe2979c20826a6d618628f839812 |
/opt/homebrew/Cellar/codec2/1.2.0/lib/libcodec2.1.2.dylib |
表示与计账边界¶
- codec 表示:headerless Codec 2 mode bitstream: 48 bits/40 ms at 1200, 48 bits/20 ms at 2400, or 64 bits/20 ms at 3200;
- 当前存储/传输 artifact:EVSC2V01 authenticated wrapper around the headerless bitstream; wrapper carries mode, original 16-kHz length, raw bitstream length/hash and frozen binary identities;
- 分账规则:representation bits, complete serialized payload bytes, program/library footprint and endpoint external checkpoint bytes remain separate; nominal mode rate never substitutes for measured payload size;
- source-blind decoder:true;
- 不能由 D0 推出的结论:deep understanding of the sinusoidal/LPC model, pitch, voicing, LSP quantization, phase or synthesis paths;the trained/optimized versus hand-designed status or scalar count of every compiled table;reproduction of any author listening test, demo comparison, PESQ/MOS or radio result;quality, intelligibility, speaker identity or expression at 1200/2400/3200 bit/s;robustness to noise, bit errors, modem/FEC/channel behavior or deployment latency;completion of the independent D5 archive despite existing parent benchmark rows。
D1 · 算法与证据深读:Codec 2 不是“小号神经 codec”¶
D1 结论:passed。 本章把作者维护的 2025 算法手册、1997 博士论文、2011 作者演讲、项目主页、文档讨论串,以及固定 release 1.2.0 的 README、
codec2.c和 CTest 入口固定为 57 条可定位主张。通过 D1 表示“知道算法为什么这样设计, 也知道官方证据缺什么”;不表示 1.2.0 已通过上游测试,也不表示作者听感结论 已被本地复现。
一句话心智模型¶
Codec 2 不尝试保留原波形或学一个通用 latent。它先做一个很强的假设:通信中的 人声,可以在每个短时窗里近似成一组以基频为间隔的正弦谐波。然后只发送:
- 基频/音高
F0; - 谐波幅度的短时频谱包络;
- voiced/unvoiced 判断;
- 帧能量。
谐波频率可由 F0, 2F0, ... 重建,相位在 encoder 丢弃、由 decoder 合成,所以不需要逐样本
波形匹配。这是它能从 8 kHz × 16 bit = 128 kbit/s 压到 1.2–3.2 kbit/s 的根本,也是其
伪影的根本:模型错了时,丢掉的波形信息无处可找。[C2-M-001–C2-M-006]
资料边界:这不是一篇可直接复现的 SOTA paper¶
Codec 2 没有像 EnCodec 那样一篇同时冻结 architecture、checkpoint、dataset 和指标表的论文, 也没有像 Opus 那样由标准组织定义的 normative decoder specification。当前 31 页手册是作者 在 2023–2025 年汇总的 living document;它要与 C 源码和自动测试一起读。作者在邮件讨论中 说得更直接:各 mode 不互操作;项目整体仍被视为可演进;他倾向用“文档 + 参考 C + 自动测试”而不是纯纸面规范来定义实际算法。[C2-B-003–C2-B-005]
因此本档案严格分三个时间层:
- 1997 lineage:博士论文提供 NLP pitch、harmonic sinusoidal analysis、phase 和 LPC/LSP 的研究起点,但其最终量化 coder 是 8.3 kbit/s,不是今天的 1200/2400/3200。
- 2011 early Codec 2:DCC 幻灯展示 2550 bit/s alpha 和当时的方向,不是 1.2.0 的 数值合同。[C2-A-011, C2-A-012]
- release 1.2.0:今天真正需要 D2/D4 验证的是 commit
06d4c11e...的函数、表、 bit packing 和 CTest;2025 手册只能提供概念地图。[C2-S-001–C2-S-008, C2-I-001]
一个已经发现的具体例子说明为什么必须分层:2025 手册 Table 2 将 2400 mode 的整帧
写为 50 bits,但同表又标 20 ms / 2400 bit/s,数学上应为 48 bits;1.2.0 源码的
codec2_bits_per_frame() 和 codec2_encode_2400() 都明确为 48,且实际分配为
36 + 8 + 2 + 2 spare = 48。所以手册很有价值,但它不能替代 release-exact code。
[C2-M-023, C2-S-002, C2-I-002]
第一性原理:为什么能压到 1.2–3.2 kbps¶
1. 发送语音生成模型,而不是波形¶
对 voiced speech,短窗内的频谱能量会集中在 F0 的整数倍上。Codec 2 把一帧写成谐波
正弦和,所有谐波共享一个 F0。对 unvoiced speech,这个模型并不真实,但 decoder
可通过随机相位去模拟噪声。这个归纳偏置丢掉了非语音、突变和复杂背景噪声的许多细节,
但把原本随时间变化的数百个 PCM 点降到少数有物理意义的参数。[C2-M-003, C2-M-004]
2. 用 LPC/LSP 压缩“哪些谐波响”¶
谐波数量 L = floor(pi/Wo) 随 pitch 改变,直接量化每个幅度会产生变长参数和很高码率。
1200–3200 modes 将谐波幅度包络近似成 10 阶 all-pole vocal-tract filter,再把 LPC 转为有序、
更适合量化的 10 个 LSP frequencies。decoder 重建 LPC spectrum,并在每个谐波频带上取能量
得到幅度。[C2-M-014–C2-M-016]
这是一种极强的“最小充分表示”尝试:不传每个谐波,而传一个能生成全部谐波幅度的 过滤器。代价是 LPC 会在高音说话人上追踪个别谐波而非包络,且低频边界、频谱峰带宽和非平坦 excitation 都会造成偏差。[C2-M-017]
3. 相位不发,在 decoder 内构造¶
相位是 Codec 2 最激进的删减。voiced frame 的激励相位按基频连续推进,vocal-tract filter 提供 minimum-phase 分量;unvoiced frame 则用随机相位。这省掉了大量 bits,但作者手册也明确 记录了代价:耳机下质量下降、voicing error 会产生 click/static,与量化后幅度结合时还会 继续降质。[C2-M-011–C2-M-013]
4. 10 ms 内核,20/40 ms 才发送¶
算法每 10 ms 做 pitch、voicing、幅度和语道分析,但为了省 bits,只每 20 或 40 ms 传某些 参数,中间帧在 decoder 线性插值。这个区别极其重要:analysis frame rate 不是 transmission frame rate。码率下降的一部分本质是降低参数时间分辨率,不只是减少量化精度。 [C2-M-005, C2-M-007, C2-S-003–C2-S-005]
编码器:从波形到少量参数¶
NLP pitch:利用非线性把“丢失的基频”变回来¶
F0 估计不能只找频谱最大峰,因为麦克风和语道会把基频削弱,而高次谐波可能更强。
NLP 先对语音平方,利用 intermodulation 在 F0 产生分量;再 DC notch、低通、降采样、DFT
找峰,并搜索全局峰的子倍频,最后用谐波对齐能量做细化。[C2-M-008, C2-A-001]
1997 论文对 NLP 的评价比一个单一 score 更值得学习:它同时使用自动 pitch-reference match、人工 contour inspection 和通过正弦 coder 听伪影。自动测试里 NLP 报告 553 hits / 50 misses / 91.7% correct,但作者发现自动分数更高不一定更好听;加 backwards tracker 后自动 正确率降到 88.1%,感知上重要的错误反而更少。[C2-A-002–C2-A-005]
这是一个早期但很深的 evaluator lesson:帧错误平均数没有按感知后果加权。高能 voiced 段的 pitch doubling 和低能 transition 中的偏差不应该同价。[C2-I-003]
Voicing:一个 bit 承担的大责任¶
Codec 2 用 MBE 变体比较实际频谱与“理想谐波系列”的拟合误差,从 1 kHz 以下能量 构造 SNR,再以经验阈值输出 binary voiced/unvoiced。单 bit 使得码率很低,但混合声源、 转换段和噪声会迫使它做过于粗糙的决策。相位合成又依赖这个决策,所以 voicing error 会被放大成 click、buzz 或粗糙背景声。[C2-M-009, C2-M-012]
Quantisation:用显式码本和标量规则代替 learned latent¶
3200 用 scalar-quantised LSP differences,2400 用 scalar LSP + joint pitch/energy VQ,1200 用 27-bit
three-stage LSP VQ + 两个有跨帧预测状态的 pitch/energy VQ index。这些 source-embedded tables 是算法状态,不能因为没有
.pt checkpoint 就宣称“零参数”。D2 将区分:手工系数、经验阈值、由训练数据生成的
codebook,以及 runtime state。[C2-S-003–C2-S-006, C2-I-004]
解码器:从参数到可懂语音¶
decoder 先拆出 voiced、pitch/energy 和 LSP,将 20/40 ms 才收到的参数插值回 10 ms。LSP
转回 LPC 后,通过每个谐波频带内的 LPC spectrum energy 恢复幅度。频域 post-filter 加强
formant peaks、降低 inter-formant energy,其 beta=0.2 与 gamma=0.5 是经验常数;作者认为它尤其
改善低音男声,但机理未完全理解,设错也会降质。[C2-M-018, C2-M-019]
相位由 excitation phase + minimum-phase LPC filter 合成,voiced 段让 pitch pulse 在帧间连续,unvoiced 段随机化。然后把谐波幅度和相位放进 DFT bins,做 IFFT 得到时域帧,用三角窗 overlap/add 拼回连续语音。[C2-M-010, C2-M-020]
这个 decoder 本质上是一个强结构生成器:它不是去找原始 waveform,而是生成一个与所收参数 一致的“合理人声”。因此感知上的自然度、speaker identity 和微妙表达可能在内容仍可懂时 已变化;这些不能由一个 waveform metric 代表。[C2-I-005]
1200、2400、3200 三个模式¶
| mode | 外层 frame | raw bits | 主要参数 | 本质取舍 |
|---|---|---|---|---|
| 3200 | 20 ms | 64 | 50-bit LSP differences + 7-bit pitch + 5-bit energy + 2 voicing | 最高时/频精度;LSP-difference error 可累积 |
| 2400 | 20 ms | 48 | 36-bit scalar LSP + 8-bit joint pitch/energy VQ + 2 voicing + 2 spare | 同样 20 ms,但用 joint VQ 省 16 bits |
| 1200 | 40 ms | 48 | 27-bit three-stage LSP VQ + 2×8-bit predictive joint pitch/energy + 4 voicing + 1 spare | 码率减半主要靠 LSP 每 40 ms 才更新 |
此表的 raw frame 长度以 release-exact C 为准。[C2-S-002–C2-S-006]
它揭示了一个对 AI 进化非常重要的事实:“low-bit 化”不是单轴 compression knob。1200 不只是 3200 的粗量化版;它同时改了参数更新节奏、量化器类型、帧延迟和 error-memory 结构。不同码率 可能需要不同的 representation,而不是强迫一个模型线性缩放。[C2-I-006]
手册还特别提醒:3200 的 LSP differences 以及 2400/1200 的 joint-delta pitch/energy 会产生误差传播, 不适合高 BER。这与手册开头“Codec 2 在几个百分点 BER 仍可懂”的宽泛描述不能无条件合并; 鲁棒性必须绑定 mode、bit importance、FEC/interleaving 和通道模型。[C2-M-021, C2-M-024, C2-I-007]
官方怎样 evaluation¶
A. 1997 论文:方法学很强,但不是当前 codec 基准¶
论文把 evaluation 分成 objective、formal subjective 和 informal subjective,并直说低码率 sinusoidal coder 丢相位后,SEGSNR 等 waveform measure 并不适用;它分别用 spectral distortion 评 LSP、用自定义 SNR 评 phase,再用听感检查伪影。[C2-A-006]
论文附录公开了它的 informal protocol:总共 196 秒,包含 2 位澳洲男声、2 位澳洲女声、 25 位英国说话人和 BBC 快速新闻;作者以约 3 秒片段对比,用耳机和 hi-fi speaker,很多判断 由作者一人完成。这是诚实、可理解的开发测试,但没有随机化、盲测、多听者原始票和置信区间。 [C2-A-007]
LSP quantiser 的 1997 数据反而特别有启发:22 分钟 TIMIT 训练,24 秒小测试集上 average SD 1.05 dB,但到 120 秒 BBC 新闻上变成 2.31 dB,49.45% 帧超过 2 dB、13.76% 超过 4 dB。 作者明确将它归因于训练材料太少、录音和频谱条件不同。这是“小数据 codebook 域外崩溃”的 早期清晰证据,但是该 8.3 kbit/s 原型的结果,不是 Codec 2 1200 的官方 SD。[C2-A-008, C2-A-009]
B. 2011 DCC:算法路线快照,没有评分协议¶
DCC 幻灯已经包含 pitch、MBE voicing、LPC/LSP、energy、post-filter 和 phase synthesis 的骨架, 但仅显示 2550 bit/s alpha 和 future-work 清单;它没有 dataset manifest、listener protocol、MOS/PESQ 表或可运行 acceptance threshold。它可以证明概念沿革,不能作为 1.2.0 的论文数字进行交叉复现。[C2-A-011, C2-A-012]
C. 当前作者手册/主页:有定性结论和样音,缺正式得分资产¶
作者主页提供原声、Codec 2 modes 和若干对照的听感链接,手册也做了“3200 在同码率 可与闭源 codec 比较”、“1300 可用于 HF”等定性陈述。但当前固定资产中没有与 1200/2400/3200 一一对应的现代标准数据集、baseline binary identity、raw votes、formal listening protocol 和可验证统计。 所以 D4 可以审计与播放作者样音,不能伪造一个本地 MOS 来宣称论文复现。[C2-B-001, C2-B-002]
手册对 700C 的德语和日语问题有作者经验性记录,同时说明 VQ 仅用 120 秒训练数据。 这是重要 risk signal,但它既不是选定的 1200/2400/3200,也没有给出德/日语数据集、人数或分数; 不能扩展成“Codec 2 不支持德语/日语”。[C2-M-022, C2-I-008]
D. release 1.2.0 CTest:可运行,但主要是 smoke/construction test¶
release CMake 为 3200、2400、1200 各有一条 c2enc | c2dec | sox 测试。它们能验证进程成功、
产生 WAV,但没有对这三条设置 reference hash、sample tolerance 或感知阈值。相比之下,700C
还有 Octave-port internal-state comparison,但它不属于本次三个选定点。[C2-S-007, C2-S-008]
因此 Codec 2 的“官方 evaluation”应该被分解为:
| 证据家族 | 能复现什么 | 不能推出什么 |
|---|---|---|
| release API/bit asserts | frame/bits/mode contract | 音质 |
| upstream CTest | build/runtime/smoke,部分 internal port | 论文 MOS 或跨实现 bit-exact |
| 1997 objective tables | 早期组件与域外风险 | 当前 1.2.0 mode score |
| author samples/qualitative notes | 作者的听感边界 | 盲测统计、多语通用性 |
| EvoSpeech D5 | 同语料/同指标/真实 bits 横评 | 上游论文原始协议的代替 |
证明了什么、没证明什么¶
已有较强一手支持¶
- Codec 2 的主干是 harmonic sinusoidal model,选定三 mode 用 LPC/LSP 表示频谱幅度;
- NLP pitch、MBE-derived binary voicing、phase synthesis、IFFT overlap/add 和 LPC post-filter 有作者数学手册;
- 1200/2400/3200 的 40/20/20 ms、48/48/64 bits 和 bit allocation 有 release-exact C 源码;
- 1997 研究已清楚识别客观指标与听感不一致、小训练集域外失败等问题;
- 当前 upstream 明确将 C 源码和自动测试视为算法定义的一部分。
尚未建立¶
- 任一选定 mode 在标准公开语料上的作者 MOS/PESQ/POLQA/ViSQOL 目标值;
- release 1.2.0 与作者网页所有历史样音的 binary/protocol 一致性;
- 德、法、西、中、日、韩的词义保真、实体正确率、说话人或表达保留;
- 选定三 mode 在真实 BER/FEC/modem/radio 下的鲁棒性;
- “与同码率闭源 codec 相当”的现代、可审计多听者证据;
- 在 2026 神经 codec 的任何给定 latency/compute/rate slice 上仍是 SOTA。
正确定位是:Codec 2 是一个极低码率、低资源、可解释的传统参数 codec 锚点。它的价值不是 预先获得“音质冠军”,而是迫使我们回答:一个更大的 SOTA 模型在 1.2–3.2 kbps 时,它多花的算力和 参数到底买到了什么?是更好的内容、identity、expression、noise generalisation,还是只有更好的 waveform metric?
对 EvoSpeech 的启发¶
- 明确写出 inductive bias。 Codec 2 因为直说“人声=谐波+语道包络+声带状态”,每个错误 都能对应 pitch、voicing、LSP、phase 或插值。未来神经 baseline 也应该把 latent 中的假设说清。
- 不要迷信单一 aggregate。 NLP tracker 的客观正确率变差但更好听,证明 evaluator 必须对错误 后果加权,并保留 breaker examples。
- “降码率”可以改变表示家族。 3200/2400/1200 不是只改一个码本数;自动科研应允许 frame rate、quantiser、prediction 和 mode routing 一起变化。
- 小 codebook 的核心是 domain coverage。 1997 LSP 域外崩溃和 700C 训练不足都提醒我们:多语、麦克风、 噪声和说话人覆盖不是最后加一个 test set,而是 representation 能否泛化的一部分。
- 把参数体积和运行代价分开。 没有 checkpoint 不等于没有训练 codebook;静态表、代码、RAM、 encoder compute 和 decoder compute 要分开计账。
- 把感知损伤归因到 mechanism。 click、buzz、reverberation、rough noise 可以生成定向反例; AI 进化的 breaker 不应只是随机新句子。
- 把 codec 与 radio stack 分开。 Codec 2 raw bits、wrapper、FEC、modem、interleaver 和 RF channel 是不同层; 不能用 clean round-trip 宣称完成 HF/VHF 通信评估。
- SOTA-first 与 low-bit 锚点不冲突。 Codec 2 负责提供可解释的下界和失败类型;现代 SOTA 负责检验更强表示在同 bits 下能否“降维打击”。两者都是 Pareto 地图必需的锚点。
D1 形成的 D3/D4 入口¶
| claim family | D1 判定 | D3 必须冻结 | 预期 verdict |
|---|---|---|---|
| 1.2.0 frame/bit allocation | release C 完整 | API query、bit count、all-zero/impulse/random 边界 | exact |
| selected-mode upstream CTest | 命令完整,acceptance 弱 | fixed source build、CTest names、exit/output/hash | exact smoke |
| deterministic encode/decode | 源码 asserts 存在 | repeated-run bitstream/output hashes、source-blind decode | exact project extension |
| 700C Octave internal port | 完整但不在 D5 点 | 作为 upstream-suite coverage,不冒充选定点 | exact auxiliary |
| 1997 pitch results | 数字/方法在论文,原始 DB 未冻结 | asset、reference labels、historical code audit | likely blocked |
| 1997 subjective conclusions | informal/small/single-author | 样音可用性、票和 protocol audit | blocked as formal replay |
| author website samples | 可听但无统一 protocol | URL/hash/mode/baseline/binary 对应 | audit-only |
| 1200/2400/3200 public benchmark | 项目已有父证据 | 180 rows、raw bits/payload、quality/content/resource | exact project protocol |
| BER/FEC/radio | 属于另一层 | mode-specific bit errors、FEC、interleaving、channel | D6/D7 独立协议 |
D2 不会再重复概括手册,而是进入 fixed source 06d4c11e...:从 codec2_create() 的 mode dispatch,
跟到 analyse_one_frame()、NLP/voicing/LPC/LSP quantisers、每个 selected encode/decode function、bit packing,再到
phase/post-filter/IFFT synthesis。同时计数并分类所有静态表,只有这样才能把 paper-to-code map 和运行合同真正
对上。
Primary-source artifact manifest¶
| Source | pages / bytes | SHA-256 | 用途 |
|---|---|---|---|
| Codec 2 Algorithm Description | 31 pp / 316,470 B | 957f9b8163c4bd34b573593c035bcee79204450a373f027377c36a48e1050f49 |
living algorithm map |
| Techniques for Harmonic Sinusoidal Coding | 155 pp / 829,457 B | bdef92ddacf9431613bef0c0a7ede0f883e0bb90415ddac40da528a284ea7585 |
NLP/sinusoidal/evaluation lineage |
| DCC 2011 Codec 2 slides | 19 pp / 461,077 B | 7610160ea25b4c1c95165ef37e0fe68c8f08949ee622b878d2abf3a765fd541c |
early Codec 2 architecture snapshot |
| Author project page | 24,778 B | e8765e9f9847c8fcaedebcaccffbd3ba8b7aa69275ffa4d6253b5719c95ade9e |
purpose and sample boundary |
| Algorithm-document discussion | 88,106 B | ac0b743ba9aa958502ca1f305b1529343ea806c5cf6f9531864fbd325b147c8c |
author scope/interoperability statements |
| Release 1.2.0 README | 9,268 B | da7db4e687ba20ce687e8e7187f59bb29f95c40c303806b710fd97027d924909 |
build/CLI/test boundary |
| Release 1.2.0 codec2.c | 61,415 B | de5b4f46a2081b3fc8d20edf119df6154c69bb7f8117bab61b841167be8972f3 |
release-exact mode contract |
| Release 1.2.0 CMakeLists | 76,579 B | d67be464eb18a0664c91c4197de00ab92e85064528ac1cfa11bd045ae74ede41 |
official test inventory |
2025 手册和 2011 幻灯已全部逐页渲染检查;1997 论文则渲染检查了目录、speech-quality
method、NLP evaluation、phase listening、quantiser test 与附录听感协议等 20 个关键页。文本抽取只用于
定位和交叉核对,没有代替图表、页码和脚注的视觉审查。结构化逐条证据见
research/models/codec2/paper_claims.json。
Codec 2 1.2.0:核心代码深读与确定性执行证据¶
D2 · 核心代码深读¶
本页回答的不是“Codec 2 听起来好不好”,而是四个更靠前、可以由源码与执行直接裁决的问题:
- 固定的 1.2.0 release 到底执行了什么函数链;
- 1200、2400、3200 三档的差别到底在哪里;
- 没有神经网络 checkpoint 是否真的等于“零参数”;
- evaluator 收到的 bytes 究竟是表示本身、上游可选文件头,还是 EvoSpeech 为 source-blind decode 加的传输元数据。
D2 使用的唯一上游代码身份是 tag 1.2.0、commit
06d4c11e699b0351765f10398abb4f663a984f36、release archive SHA-256
cbccae52b2c2ecc5d2757e407da567eb681241ff8dadce39d779a7219dbcf449。42 个实际阅读和计账的源文件分别写入
code_map.json;函数行号也全部指向该 commit。运行证据来自本机 Homebrew 1.2.0 arm64 bottle 中固定 hash 的
c2enc、c2dec 与 libcodec2.1.2.dylib。
这里必须先订正 D1 的一个术语。源码常量叫 LSP_PRED_VQ_INDEXES,但 1200 路径的
encode_lsps_vq() 实现是一个三级 LSP VQ:第一阶段量化完整十维向量,第二、三阶段分别量化偶数和奇数位置的残差。
它本身没有跨帧 LSP predictor。真正有跨帧预测状态的是 2400/1200 共用的 joint pitch-energy VQ:
xq_enc[2] / xq_dec[2] 经过 0.8、0.9 两个 predictor coefficient 递推。D1 文档和结构化 claim 已同步改为
“three-stage LSP VQ + predictive joint pitch-energy VQ”,避免从名字反推不存在的机制。
固定代码面与可信边界¶
Codec 2 的可执行定义不是一个文件。最小闭环由以下几层构成:
| 层 | 固定文件 | 本轮读取的事实 |
|---|---|---|
| public API | codec2.h |
mode 常量、create/destroy、encode/decode、bits/samples per frame |
| route/state | codec2.c, codec2_internal.h |
mode dispatch、共享分析/合成、跨帧状态、三档实现 |
| analysis | nlp.c, sine.c |
非线性 pitch、谐波 refinement、幅度、MBE voicing |
| envelope | lpc.c, lsp.c, quantise.c |
LPC/LSP、能量、标量/VQ、错误恢复与 postfilter 输入 |
| decoder | interp.c, phase.c, postfilter.c, sine.c |
插值、phase generation、背景噪声处理、IFFT/overlap-add |
| representation | pack.c, c2enc.c, c2dec.c |
MSB-first fields、raw stream、可选 .c2 header、CLI 边界 |
| build tables | src/CMakeLists.txt, generate_codebook.c, src/codebook/*.txt |
哪些表进入 binary、表维度与生成方式 |
| project transport | codecs_/codec2_payload.py |
.bin raw representation 与 EVSC2V01 wrapper 分账 |
这也是为什么“读论文摘要 + 跑命令”不算完成。手册解释为什么使用正弦模型,codec2.c 决定当前模式每一位的
含义,codebook 文件决定运行时搜索空间,CLI 文件又决定输出后缀是否偷偷加 header。任何一层没读,都可能得到一个
数值看似合理、身份却错误的 benchmark。
本轮没有把整个 FreeDV repository 都算作 Codec 2 endpoint。调制、FEC、OFDM、FreeDV framing、700C/newamp、 LPCNet/FreeDV 2020 都存在于同一源码树,但不在当前 1200/2400/3200 的干净文件路径中。它们以后若作为新的 endpoint, 必须重新固定 build flag、模式、码流和参数计账,不能借用本页结论。
状态、初始化与 mode dispatch¶
codec2_create(mode) 先检查 mode 是否在编译时启用,然后分配一个 struct CODEC2。这个对象不是一个无状态函数的
包装;它同时容纳 encoder 和 decoder 的长短期记忆:
- 8 kHz、10 ms hop、pitch window 等派生常量;
- analysis/synthesis FFT plan 与 window;
- 输入 speech history
Sn、输出 overlap/add historySn_; - NLP estimator 的滤波、平方信号与 previous-F0 state;
- previous decoded sinusoidal model、previous LSP、previous energy;
- joint pitch-energy VQ 的 encoder/decoder 两组二元素递推状态;
- excitation phase、background estimate、LPC postfilter 的 beta/gamma 与 bass-boost 开关。
初始化随后把 encode / decode function pointer 绑定到 mode-specific 函数。对当前三档分别是
codec2_encode_3200 / codec2_decode_3200、...2400、...1200。公共
codec2_encode() 不再判断码率,而是调用已绑定函数。这种结构的意义不只是 C 工程风格:它说明 mode 是 decoder
状态身份的一部分,不能从一个 headerless byte array 可靠推断。
codec2_bits_per_frame() 与 codec2_samples_per_frame() 给出 release-exact 合同:
| mode | samples/frame | bits/frame | external frame | nominal rate |
|---|---|---|---|---|
| 1200 | 320 | 48 | 40 ms | 1.2 kbit/s |
| 2400 | 160 | 48 | 20 ms | 2.4 kbit/s |
| 3200 | 160 | 64 | 20 ms | 3.2 kbit/s |
内部 n_samp 固定是 80 samples,即 10 ms。外部 frame 只决定一次调用聚合两个还是四个 internal hop,以及哪些参数
在哪一个 hop 才更新。这是理解 low-bit 的关键:bitrate 不只决定每个数用几位,还决定时间采样结构。
真实 encoder 调用链¶
三档共享的 encoder 骨架可以压缩为:
PCM16 frame
→ 每 80 samples 调 analyse_one_frame
→ shift Sn history + dft_speech
→ nlp coarse pitch
→ two_stage_pitch_refinement
→ estimate_amplitudes
→ est_voicing_mbe
→ 在指定 hop 调 speech_to_uq_lsps
→ autocorrelation → Levinson-Durbin → LPC energy
→ LPC bandwidth expansion → LPC-to-LSP roots
→ mode-specific pitch/energy/LSP quantiser
→ pack(MSB first, exact field width)
→ assert packed bit count == mode contract
analyse_one_frame() 每次把旧 Sn 左移 80 个样本,放入新 speech,计算 windowed DFT,再依次调用 pitch、harmonic、
amplitude 和 voicing 模块。它只形成未量化的声源/谐波模型;LPC/LSP 并不在这里每 hop 无条件执行,而由各 mode 在需要
发送 spectral envelope 的 hop 上调用。这解释了为何 1200 第 2 个 hop 仍会做一次 LPC energy 分析,但直到第 4 个 hop
才发送 LSP:energy 与 spectral-envelope update schedule 并不完全相同。
编码函数在打包前把目标 buffer 清零,并在最后 assert(nbit == codec2_bits_per_frame(c2))。这条 assert 能抓到开发者
改变 field width 却忘记更新 frame contract 的错误,但不是互操作 conformance:它没有独立 decoder、golden bits 或
语音质量 oracle。
NLP pitch、harmonic refinement 与 voicing¶
nlp() 的“non-linear”不是神经网络。它把最近的新 speech samples 平方,经 DC notch 与 low-pass FIR,按固定比例
decimate 后送入 DFT。然后在允许的 pitch-period 范围搜索 global spectral peak,再用
post_process_sub_multiples() 检查整数 submultiple。这个 postprocessor 会参考 previous F0,并包含作者在源码中明确称为
experimentally derived / “magic number”的 threshold。最后得到的 coarse period 还不是最终 Wo。
two_stage_pitch_refinement() 在 coarse 周围做两级 harmonic-sum search,让 integer harmonics 与 speech spectrum 更好
对齐。estimate_amplitudes() 在每个谐波 band 内累计 spectral energy,形成 MODEL.A[m];正常 selected path 不需要
保存原始 phase。est_voicing_mbe() 比较观测 low-frequency spectrum 与理想 harmonic reconstruction 的误差,再用
threshold 得到一位 voiced/unvoiced。
这条调用链解释了 D1 中历史 pitch 结果为何不能直接当 codec quality:pitch tracker 可以降低 frame accuracy 却减少更 难听的 gross error;voicing mistake 又会改变 phase generation,从而产生 click/static。一个 upstream 单元的分类误差, 会经生成式 decoder 非线性放大,最终损失取决于能量、语音位置和相邻状态,不等价于均匀的 frame error count。
同样,本轮的确定性多音 probe 不是 pitch benchmark。它只迫使 analysis/synthesis 路径实际执行并检查 framing;没有 pitch ground truth,因此绝不产生“pitch estimator 正确”的结论。
LPC/LSP:把谐波幅度压缩成十个频率¶
speech_to_uq_lsps() 对 windowed speech 做 autocorrelation,使用 Levinson-Durbin 得到 10 阶 LPC,另外计算 residual
energy。它先基于未扩展 LPC 计算 energy,再给系数施加 15 Hz bandwidth expansion,降低 LSP root finding 偶发失败的
风险。lpc_to_lsp() 若没找到恰好十个 root,就返回一个均匀、良性的 LSP vector,而不是让无效 spectrum 继续传播。
LSP 比直接量化 LPC 更适合这里,因为它有有序频率的解释并可较安全插值,但代码仍需要两类保护:
check_lsp_order()修正解码后失序;bw_expand_lsps()强制低频/高频段的最小间距。
这些保护也是“decoder 生成 plausible waveform”的证据:收到的离散 index 不会机械还原原波形,而是进入一个带稳定性 约束、插值、修正、postfilter 的语音生成器。它可能保持内容但改变 speaker/expression;D2 不能用结构合理性代替 D5 的 实测。
aks_to_M2() 再将重建 LPC spectrum 映射回每个 harmonic magnitude,并可执行 LPC postfilter、bass boost。这样传输
十个 LSP + energy 就能生成数量随 pitch 变化的 harmonic amplitudes,这是从 waveform 维数“降到最小声学描述”的核心。
3200、2400、1200:不是同一旋钮的三个刻度¶
3200 的外部 20-ms frame 包含两个 10-ms analysis hop。第一个只发 voicing;第二个发 voicing、7-bit linear pitch、
5-bit scalar energy 和 50-bit LSP differences。encode_lspds_scalar() 不是把十个绝对 LSP 独立量化:第 i 个值是当前 LSP
与“前一个已经量化的 LSP”之差,因此一个错误可影响同一帧后续频率重建。十个 difference grid 各 32 levels,共 320 个
设计型 scalar values。
2400 保持 20-ms frame 与两个 10-ms voicing decisions,但做两次架构替换:
- pitch 与 energy 不再分别用 7+5 bits,而用一个 8-bit joint VQ index;
- LSP 改为十个 absolute scalar codebooks,总计 36 bits,而非 50-bit differences。
再加两个 spare bits,恰好 48。joint VQ 在 log-pitch/log-energy domain 中计算感知加权残差;此前重建状态分别乘 0.8、
0.9 后参与预测。它节省位数,却要求 encoder/decoder 的 xq history 完全同步。一个坏 index 的影响不一定在当前 frame
结束,因此 clean-file trace 与 BER robustness 是两个不同实验。
1200 每次吃 40 ms,也就是四个 analysis hop。voicing 每 10 ms 仍发一位;joint pitch-energy 在第 2 和第 4 hop 各发 8 bits;LSP 只在第 4 hop 更新一次。十维 LSP 的第一级有 512 个中心(10×512 floats),随后偶数与奇数坐标的 residual 各用一个 5×512 codebook,三个 index 各 9 bits,总计 27。再加 1 spare bit,整帧 48 bits。
因此从 3200 到 1200 同时改变:external frame duration、LSP update rate、quantiser family、是否使用 cross-frame pitch-energy prediction、table footprint、interpolation span 与 error-memory。对 EvoSpeech 的直接启示是:未来自动进化不应 只搜索“RVQ stage 数”或“token rate”;更新频率、状态同步、参数耦合和 decoder prior 都应是可改变的研究变量。
decoder、插值、phase 与 overlap/add¶
selected decoder 先按 mode 的 exact field order unpack()。没有发送的 10-ms 值由前一 frame 状态和本帧 endpoint
interpolate:pitch 的插值还会根据前后 voiced flag 选择 midpoint、前值、后值或 Wo_min;energy 用 geometric mean;LSP
则按 0.5(20 ms)或 0.25/0.5/0.75(40 ms)线性插值。
每个恢复的 10-ms model 随后执行:
LSP → lsp_to_lpc → aks_to_M2 → apply_lpc_correction
→ sample_phase → phase_synth_zero_order
→ postfilter → synthesise(IFFT + overlap/add)
→ ear_protection → PCM16 clip
sample_phase() 从 reconstructed LPC filter 取得每个 harmonic 的 synthesis-filter phase。
phase_synth_zero_order() 对 voiced frame 延续 ex_phase += Wo*n_samp,对每个 harmonic 乘以 harmonic number;对 unvoiced
frame 则用 codec2_rand() 产生随机相位。源码不传 pulse position,也不传原始 phase trajectory。
这里出现了一个实现层面的并发风险:codec2_rand() 的 LCG state next 是 sine.c 文件级 static,不是
struct CODEC2 成员。两个 decoder instance 在同一进程交错执行时会共享这个全局 sequence;它也没有显式线程同步。
本轮 CLI trace 每次启动新进程,因此重复 decode 从相同初值开始并逐字节一致,但这不能证明一个多实例、并发 library
deployment 具有同样的 instance-local determinism。这个差距已进入 release-gap ledger,而不是被重复性数字掩盖。
decoder postfilter 还维护 bg_est:只在低能量 unvoiced frame 更新背景估计,在 voiced frame 中把低于背景 margin 的
harmonic phase 随机化,试图降低 clicky background noise。这不需要发送额外 bits,但把 history、threshold 与 RNG 变成输出
的一部分。最后 synthesise() 把 complex harmonics 放进共轭频谱,IFFT 后用 trapezoid window 与旧 Sn_ overlap/add;
ear_protection() 与 PCM16 clipping 负责限制异常能量,而不是修复 bit error。
参数、表、阈值与 runtime state 的四本账¶
“传统 codec 没有 checkpoint,所以 parameters=0”是错误口径。本轮按来源与角色拆成四类:
- data-derived VQ table:1200 的三阶段 LSP VQ 共有 10,240 floats,2400/1200 共用的 pitch-energy
gecb有 512 floats,unique total = 10,752 parameter elements。它们决定 nearest-neighbour search 的中心,应进入 learned / data-derived endpoint parameter count。 - designed scalar grid:2400 absolute-LSP grids 132 floats,3200 difference-LSP grids 320 floats,合计 452。文件中 是规则化、可直接审阅的量化 levels;本页不把它们伪装为 corpus-trained weights,但仍属于 algorithm tables。
- designed coefficients / empirical thresholds:如 Wo/E predictor
[0.8, 0.9]、LPC postfilter beta/gamma、NLP FIR、 voicing threshold、background threshold/margin。它们影响行为,但和从数据得到的 VQ centres 分账。 - runtime working state:FFT plan/window、speech buffer、NLP history、previous F0/model/LSP/energy、encoder/decoder xq、 excitation phase、overlap/add、background estimate。它随 stream 变化,是 memory/state,不是 learned parameter。
10,752 的边界是“三个 selected modes 可到达的 unique VQ floats”,不是整个 libcodec2 二进制里所有 mode 的 table。
700C/newamp tables 虽然也被默认 library build 编译,当前 endpoint 不执行;installed dylib size 也另行报告,不能把 file bytes
除以四冒充 parameter count。
另一个限制是训练 provenance。release 保留最终 lspjmv*.txt 与 gecb.txt,并在源码明确指出 joint Wo/E VQ 针对
8 kHz trained;但这个 archive 没有提供可从原始 corpus、seed、命令完整重生全部 10,752 数值的流水线。我们可以验证 table
维度、hash、生成到 C 的过程和 runtime lookup,不能声称“重训复现”。D3 会据此把 author-native reproduction 分成
可严格执行与 asset 缺失两类。
raw bitstream、.c2 header 与 EVSC2V01¶
pack() / unpack() 只操作 MSB-first fixed fields。headerless raw stream 没有 magic、mode、sample count 或 checksum。
尤其 1200 与 2400 都是六 bytes/frame,却分别表示 40 ms 和 20 ms;只看文件长度不可能知道正确 decoder。
上游 CLI 有一个容易漏掉的 suffix contract:若 c2enc 输出文件名以 .c2 结尾,它会先写一个 struct c2_header,内容
包括 magic、release version、mode 与 flags。输出 .bin 则不写 header。D2 trace 故意使用 .bin,因此 60/120/160
bytes 是纯 representation;不是把一个带 metadata 的文件大小当 nominal codec payload。
EvoSpeech 正式 benchmark 又需要 decoder source-blind 且可独立运行,所以 codec2_payload.py 在 raw stream 外封装
EVSC2V01,包含 mode、原始 16-kHz length、raw byte count/hash、固定 encoder/decoder hash。计账同时保留:
representation_bits = len(raw_bitstream)*8;raw_bitstream_bytes;- complete serialized wrapper bytes;
- runtime/library footprint 与 external checkpoint bytes。
wrapper 不是免费,也不是 Codec 2 算法位率;它是 evaluator transport metadata。formal adapter 在临时文件中使用 .bin。
D2 还发现 legacy Codec2Codec 曾使用 out.c2 却在 docstring 声称 no container,导致 payload_bytes 混入 header。本轮已
改为 out.bin 并加测试,防止旧 registry 入口与正式 benchmark 的口径再次分裂。
确定性 mechanism trace¶
scripts/trace_codec2_code.py 生成 0.4 秒、3,200 samples、8 kHz mono PCM16 的三分量确定性波形。它不是 speech sample,
不含 transcript、speaker 或 quality ground truth。0.4 秒同时整除 20 ms 与 40 ms,避免末帧 truncation 混入机制检查。
每档执行两次独立 encode,并对同一 bitstream 启动两次独立 decode。实测为:
| mode | raw frames | bytes/frame | total raw bytes | repeated encode | decoded samples | repeated decode |
|---|---|---|---|---|---|---|
| 1200 | 10 | 6 | 60 | byte-identical | 3,200 | byte-identical |
| 2400 | 20 | 6 | 120 | byte-identical | 3,200 | byte-identical |
| 3200 | 20 | 8 | 160 | byte-identical | 3,200 | byte-identical |
trace 在运行前重新 hash 42 个固定 source 和三个 runtime artifacts;任何 source、binary 或 table 改变都会中止。decoder
命令只得到 mode、raw stream、固定 binary 与输出路径,没有 source PCM。生成的每个 bitstream/decoded PCM hash 记录在
measurements/code_trace.json,并可用 python3 scripts/trace_codec2_code.py --check 重跑比较。
这里的 “packet_summary” 只是沿用项目统一 validator schema,字段明确标注为 headerless fixed-size Codec 2 frames,绝不 声称它们是 RTP/network packets。raw framing overhead 在这个 CLI 文件实验中为零;实际网络若增加 FEC、interleaver、 modem 或 packet header,必须另计。
release surface gaps:源码没有替我们证明什么¶
D2 明确保留七个边界:
- living manual 比 1.2.0 新,2400 table 还存在 50/48-bit discrepancy;exact 数字由 release source 裁决。
- raw stream 不自描述;mode metadata 是正确 decode 的必要条件,但须与 representation 分账。
- selected-mode CTests 只有 process-success,没有 independent implementation、golden vector 或 perceptual threshold。
- final VQ tables 存在,但完整 corpus-to-table training provenance 不在 selected archive 中。
- clean trace 没执行 BER、soft decision、FEC、modem 或 radio channel;stateful VQ 的 error propagation 仍待独立协议。
- 单机 byte repeat 不是 quality,也不保证跨 compiler/libm/architecture bit identity。
- unvoiced RNG 是进程级 static global state;CLI 单实例重复性不代表同进程多 decoder/thread 的 instance-local determinism。
这些不是对 Codec 2 的苛责,而是把“source exists”“program runs”“implementation conforms”“speech is good”四种不同主张 拆开。D2 只把前两种推进到强证据;第三种取决于 D3 的 official assets,第四种必须由 D5 common benchmark 和 D6 人工校准 回答。
D2 对 D3/D4/D5 的直接约束¶
D3 的 paper-native protocol 不能凭空制造一个“官方 MOS 表”。它应首先盘点 release/source 是否存在:golden bitstream、 reference PCM、independent decoder、pitch labels、LSP training/test vectors、作者 listening stimuli 与 raw votes。不同 claim 分别标为 strict、approximate 或 blocked,不能用现有 public common benchmark 反向冒充论文复现。
D4 若执行 upstream selected-mode tests,结论只能是 fixed release smoke path passed。若重跑 thesis pitch/LSP component metrics,必须固定对应历史 code/data identity,不能把 1997 result 贴到 2023 release。VQ training provenance 缺失时,应报告 “runtime table identity reproduced, training regeneration blocked”。
D5 则固定使用真正的 .bin representation 与 EVSC2V01 transport,至少报告:measured raw bitrate、complete payload bytes、
8→16 kHz normalization、narrowband-aware quality、内容、speaker、资源与公开样本 blind listening。Whisper/SenseVoice 已在父级
公共协议中执行过,不因 D2 再跑;D2 也没有恢复任何个人录音硬要求。
D2 允许与禁止的结论¶
现在可以说:
- 已把手册中的 harmonic/LPC/LSP/pitch/voicing/phase 机制逐项定位到 fixed 1.2.0 source;
- 三档确实是不同 temporal/quantiser/state architecture,而非同一算法的码率参数;
- 三档 raw frame contract 在 fixed runtime 上执行成立且对 deterministic probe byte-repeat;
- selected endpoint 含 10,752 个 data-derived VQ parameter floats、452 个 designed scalar-grid floats,并有大量独立 runtime state;
- formal project decoder 不接触 source waveform,mode/length/hash metadata 与 representation bytes 分账;
- 找到并修复了 legacy
.c2suffix 导致的 header/payload 口径错误。
仍然不可以说:
- Codec 2 的 1200/2400/3200 quality、intelligibility、speaker 或 expression 已由 D2 证明;
- 当前实现重现了 1997 thesis、2011 alpha 或作者主页听感结论;
- final VQ table 的训练过程已复现;
- fixed CLI 与任何独立 implementation bit-exact interoperable;
- clean-file repeatability证明 BER、packet loss、radio 或多线程部署可靠;
- 没有外部 checkpoint 就等于算法没有 data-derived parameters。
D2 的最终价值不是多一张漂亮架构图,而是让下一步评价不再研究一个身份模糊的“Codec 2”。从现在开始,每个指标都必须 对应:明确 mode、明确 raw bytes、明确 state reset、明确 release、明确数据与明确允许的结论。
D3 · 论文原生复现协议¶
D3 结论:passed。 协议
codec2_1_2_0_paper_native_v1已在任何正式 D4 结果产生前冻结。 Codec 2 1.2.0 的 release archive 能支持 release identity、官方命令链 smoke 和跨 build interoperability 检查;它不能支持 selected-mode 感知质量、历史 pitch/LSP 数字、听测或 VQ 训练过程的严格重放。D4 必须逐类报告passed/failed/blocked/out_of_scope,不能用新的代理实验 填补旧实验缺失的原始资产。
最重要的发现:官方 release 有 smoke assets,没有 selected-mode quality oracle¶
固定的 1.2.0 archive 包含六个 8-kHz PCM16 raw 文件、c2enc、c2dec 源码以及 CMake/CTest
命令。这足以回答一个有限但重要的问题:固定源码是否可以构建,1200、2400、3200 三条官方
encode → decode 命令是否成功,headerless bitstream 是否具有预期 frame/byte 数,decoder 是否
恢复精确的样本数。
它没有附带三档模式的 golden encoded stream,也没有 reference decoded PCM、允许误差、MOS、
MUSHRA、ViSQOL、PESQ 或 listener-level votes。上游三个 selected-mode CTest 的实质 oracle 是命令
成功;其中的 sox 只负责格式转换/输出,不比较质量。因此“CTest 通过”只能称为 executable smoke
contract,不能写成音质、可懂度或论文 headline 指标复现。
这个区分对传统 codec 尤其重要。Codec 2 的 nominal rate 是算法 bit allocation,不自动证明
implementation conformance,更不证明主观质量。D4 同时记录 raw representation bytes、.c2 header
和项目 wrapper,禁止把容器/元数据混入或移出码率口径来制造优势。
Track A · release asset identity¶
D4 运行前先验证 release tarball 的 7,688,549 bytes 与 SHA-256,并逐项验证全部六个 raw/*.raw
成员及六个固定 harness/source 文件。六个 raw 文件全部入 manifest,不只登记最终选择的
hts1a.raw;这样可以防止以后同名文件被替换、裁剪或重新编码。
这个 track 是 strict_primary_artifact_check_not_headline_metric。任何 byte/hash 不同都会停止 D4,
但全部相同也只证明输入与代码身份正确。它不会因为“官方仓库里有这个文件”就把该文件误称为
独立 conformance vector。
Track B · fixed-source selected-mode smoke¶
主执行轨在 clean Docker Linux/arm64 中从固定 release archive 构建。Docker 的作用是记录 Linux、
compiler、CMake options、linked libraries 和 binary hashes,不是把容器本身当作科学正确性的证明。
输入固定为 raw/hts1a.raw:mono signed PCM16 little-endian、8 kHz、24,000 samples、3.0 秒。
三档预注册结果只来自公开 frame contract,而不是运行后挑选:
| mode | samples/frame | bits/frame | frames | raw bytes | decoded samples |
|---|---|---|---|---|---|
| 1200 | 320 | 48 | 75 | 450 | 24,000 |
| 2400 | 160 | 48 | 150 | 900 | 24,000 |
| 3200 | 160 | 64 | 150 | 1,200 | 24,000 |
每档必须同时满足上游 CTest exit 0、直接 headerless encode 的精确 frame/byte 数、固定源码 decoder
exit 0 和精确 24,000 个输出样本。缺任一行或任一命令失败,都使 PN-C2-SELECTED-MODE-SMOKE
失败;不能以其余两档平均通过掩盖。
Track C · .c2 header 与 raw representation 分账¶
上游 test_codec2_mode_dot_c2 专门检查以 .c2 suffix 输出时写入的 file header,以及 decoder
从 header 检出 mode 8 的行为。D4 会运行该测试并记录 header bytes,但项目三档 benchmark 使用的
representation 是 headerless stream。.c2 测试通过不改变 450/900/1,200-byte raw payload,反之亦然。
这一轨直接吸收 D2 发现的 legacy adapter 问题:文件 suffix 会改变上游 CLI 行为,所以“临时文件名”
不是无关实现细节。当前 adapter 已固定使用 .bin,EVSC2V01 transport 中 mode、原始长度、hash 与
raw stream 分账。
Track D · same-revision cross-build diagnostic¶
D4 用同一个 hts1a.raw 比较两个 release-1.2.0 endpoint:Linux/arm64 fixed-source build 与
macOS/arm64 Homebrew bottle。三种 mode 分别由两个 encoder 产生六个 artifact,再交给两个 decoder,
形成 3 modes × 2 encoders × 2 decoders = 12 行。
硬通过条件只有:每个 encoder 的 raw byte 数准确,每个 cross-decode exit 0,输出都是 24,000 samples。 bitstream SHA-256 是否跨 build 相同、同一 bitstream 的 decoded PCM 是否相同、首个差异 byte/sample 在哪里,全部原样报告为平台条件诊断。因为算法包含 floating-point analysis,D3 不在看到结果前假定 不同 toolchain 必须 bit-exact;但也不允许出现差异后把它藏掉。
历史 pitch/LSP/listening 为什么 blocked¶
1997 thesis 和后续作者资料提供了有价值的算法谱系、实验方法与 aggregate 数字,例如历史 pitch tracker/NLP 百分比、LSP spectral distortion/outlier 表和听感结论。这些证据说明作者测过什么,却没有 在当前冻结资产中提供完整 hand-labelled pitch frames、逐 frame labels、exact feature matrices/splits、 historical binaries、完整 stimuli 顺序、listener-level responses、screening 与统计脚本。
因此三类 claim 在执行前就是 blocked_before_execution。blocked 不表示作者结果错误,也不表示
我们可以忽略它;它表示本项目不能计算一个有意义的 reproduction delta。换一批公共中文/英文,跑
Whisper、SenseVoice、ViSQOL 或由项目所有者听一遍,会形成“当前 1.2.0 的新实验”,不会恢复历史
实验身份。那些当前模型结果只能进入 D5。
VQ runtime identity ≠ training reproduction¶
D2 已从 release-exact 调用路径计账:selected modes 触达 10,752 个 data-derived VQ float,即按 float32 为 43,008 bytes,另有 452 个 designed scalar-grid 元素。D4 可以核对最终 table 的源码身份、 维度和 runtime 使用位置;这叫 runtime-table identity。
但 release archive 没有给出能从原始 corpus 一步生成所有最终 float 的完整闭环:缺完整训练语料/
features、预处理、split、seed、commands 和逐 table regeneration proof。因此
PN-C2-VQ-TRAINING 固定为 blocked。最终表存在不等于训练过程可复现,“传统 codec 没 checkpoint”
也不等于零 data-derived parameter。
metric、统计与失败规则¶
D4 的 strict metric 只有 artifact byte/hash、process exit、raw byte count 和 decoded sample count。 单位是确定性的 build/mode/encoder/decoder case,不从总体抽样,所以不计算置信区间。上游三档 CTest 预期 3 行,header CTest 1 行,两个 build 的 direct encoder 6 行,cross-decode 12 行。summary 报告 passed/failed/missing;missing required row 必须失败,而预分类的历史 blocker 保持 blocked。
正式输出已预注册为 machine-readable summary、逐项 JSONL、Markdown report、完整 run log 和 build identity。runner 必须先验证所有输入/source hashes;不能先跑完再选择有利文件。任何新阈值、 质量指标或数据集必须成为 D5 或新协议,不能在 D4 结果出现后回写 D3。
D3 允许与禁止的结论¶
D3 允许说:release 1.2.0 拥有足够资产执行三档 fixed-source smoke、.c2 header test、严格 byte/frame/
duration 检查和 same-revision cross-build interoperability;九个 claim family、输入身份、失败规则和
五类输出已在正式结果前冻结;历史 component/listening 与 VQ training 因关键资产缺失被明确 blocked。
D3 仍禁止说:本地 build 已通过;Linux 与 macOS bit-exact;1200/2400/3200 的感知质量已复现;历史 91.7/88.1%、LSP distortion 或听测数字已被确认或推翻;Codec 2 优于任何 neural codec;以及已有 公共 ASR/质量结果可以反向充当历史论文复现。下一步 D4 只执行本协议,完成后再单独冻结 D5 的 180-row 公共 benchmark 审计,不重跑已经完成的 Whisper/SenseVoice,也不要求个人录音。
Codec 2 D4 · paper-native reproduction report¶
Status: passed under frozen protocol codec2_1_2_0_paper_native_v1.
Official release smoke and framing¶
The release archive, all six bundled raw PCM assets and six frozen harness sources matched their D3 byte counts and SHA-256 values. The clean Linux/arm64 source build ran the three selected-mode upstream CTests and the .c2 suffix/header CTest. The two encoders produced all six expected direct streams: 450 bytes at 1200, 900 at 2400 and 1,200 at 3200, with exact 24,000-sample recovery.
| frozen item kind | expected | passed |
|---|---|---|
| upstream CTest | 4 | 4 |
| direct encode | 6 | 6 |
| cross-build decode | 12 | 12 |
This is an executable smoke/framing result, not a perceptual-quality score. The upstream tests have no golden selected-mode bitstream, reference PCM or MOS/ViSQOL threshold.
Same-revision Linux/macOS diagnostic¶
All twelve decoder/encoder combinations satisfied the hard interoperability contract. Linux and macOS encoded bitstreams were byte-exact in 1/3 modes. For the six (encoder, mode) streams, Linux and macOS decoded PCM was sample-exact in 0/6 comparisons. Any first-difference locations remain in the structured result; equality was diagnostic and was not used to rewrite the frozen smoke threshold.
Historical evaluation and VQ provenance¶
Historical pitch accuracy, LSP spectral-distortion/outlier values and listening conclusions remain blocked for strict replay because the frozen evidence set lacks exact labels/features/splits, historical executables, complete stimuli and listener-level responses. The 10,752 runtime VQ floats are identity-bound, but corpus-to-table training reproduction remains blocked. D4 generated no substitute Whisper, SenseVoice, ViSQOL or owner-listening number.
What D4 establishes¶
- fixed release source/input identity, buildability, framing, suffix/header behavior and two-build interoperability were executed under the precommitted protocol;
- process success is not reported as quality, intelligibility, identity or expression;
- current 30-English + 30-Chinese public results remain D5 evidence and are not historical-paper deltas;
- no personal recording or redundant ASR execution is required by this Gate.
Codec 2 D5 · 公共统一横评结果¶
状态:passed。这是已完成共同实验的确定性模型级归档,不是一次新的 codec、ASR 或质量指标运行。 180 个逐样本对象与父证据逐字段相等;三档 aggregate 与固定 bootstrap CI 原样保留。
三档全体结果(30 英 + 30 中)¶
| 点位 | representation bps | serialized payload bps | EVSC2V01 overhead | ViSQOL | STOI | Whisper ΔER en/zh | SenseVoice ΔER en/zh | WavLM | Resemblyzer | enc ×RT | dec ×RT |
|---|---|---|---|---|---|---|---|---|---|---|---|
codec2_1200bps |
1197.5 | 1733.3 | +44.74% | 3.7105 | 0.6302 | +5.06%/+6.42% | +3.03%/+5.00% | 0.9078 | 0.7111 | 17.66 | 21.93 |
codec2_2400bps |
2398.4 | 2934.8 | +22.37% | 3.7338 | 0.6317 | +3.18%/+2.48% | +1.09%/+2.86% | 0.9134 | 0.7181 | 18.35 | 22.29 |
codec2_3200bps |
3197.9 | 3734.4 | +16.78% | 3.7400 | 0.6416 | +2.27%/+0.07% | +0.61%/+1.03% | 0.9322 | 0.7470 | 18.00 | 23.34 |
ΔER = decoded error rate − reference-audio error rate;正值表示 codec 输出让该 ASR 更难识别。绝对数必须与 reference-audio baseline 一起解释。
语言条件与 95% CI¶
| 点位 | 语言 | ViSQOL mean [95% CI] | STOI mean [95% CI] | Whisper ΔER [95% CI] | SenseVoice ΔER [95% CI] |
|---|---|---|---|---|---|
codec2_1200bps |
en | 3.2668 [3.1478, 3.3819] | 0.6457 [0.6330, 0.6591] | +0.0506 [+0.0149, +0.0908] | +0.0303 [+0.0103, +0.0520] |
codec2_1200bps |
zh | 4.1542 [4.0148, 4.2934] | 0.6147 [0.6028, 0.6272] | +0.0642 [-0.0032, +0.1291] | +0.0500 [+0.0173, +0.0897] |
codec2_2400bps |
en | 3.2985 [3.1813, 3.4134] | 0.6470 [0.6340, 0.6613] | +0.0318 [+0.0051, +0.0648] | +0.0109 [-0.0097, +0.0397] |
codec2_2400bps |
zh | 4.1690 [4.0283, 4.3063] | 0.6164 [0.6034, 0.6300] | +0.0248 [-0.0332, +0.0808] | +0.0286 [+0.0036, +0.0588] |
codec2_3200bps |
en | 3.3274 [3.2094, 3.4381] | 0.6580 [0.6457, 0.6707] | +0.0227 [+0.0015, +0.0488] | +0.0061 [-0.0066, +0.0220] |
codec2_3200bps |
zh | 4.1526 [4.0201, 4.2819] | 0.6252 [0.6116, 0.6390] | +0.0007 [-0.0583, +0.0557] | +0.0103 [-0.0091, +0.0323] |
四本不能混的账¶
- raw representation 平均约 1.198/2.398/3.198 kbps;这是 mode frame bits 除以真实时长;
- 完整 EVSC2V01 payload 平均约 1.733/2.935/3.734 kbps,短句固定元数据使相对 overhead 为 +44.74%/+22.37%/+16.78%;
- common benchmark 的 endpoint model state 为 0,只表示没有外部 checkpoint;
- D2 仍确认所选路径含 10,752 个 data-derived VQ floats(43,008 bytes)、452 个 designed scalar-grid 元素,程序/动态库 footprint 另算。Codec 2 不能被描述为“零参数”或“零模型成本”。
读数与研究启发¶
- 1.2→2.4 kbps 的全体 ViSQOL 只从 3.7105 到 3.7338、STOI 从 0.6302 到 0.6317;3.2 kbps 才把 STOI、WavLM、Resemblyzer 和双 ASR ΔER 同时推向更好方向。这说明 Codec 2 的低档瓶颈不只是再多几个 bits,而是 8-kHz 窄带、参数表示与 mode-specific quantiser/synthesis 的共同限制。
- SI-SNR 在三档仍约为 −33 到 −30 dB,而 ViSQOL 约 3.71–3.74。参数语音会重建相位和谐波细节,不追求 sample waveform 对齐;因此未来 AI 进化若把 SI-SNR 当唯一 reward,会系统性惩罚可懂但波形不同的好方案。
- 中文 ViSQOL 约 4.15、英文约 3.27,差距在三档都存在;它更可能反映两个语料分布、录音条件和 evaluator 行为,而不是‘Codec 2 天生更适合中文’。报告必须保留分语言表,不能只看 60 句平均。
- 1.2 kbps 下 Whisper/SenseVoice 的 en/zh ΔER 都明显为正;提高到 3.2 kbps 后中文增量接近零、英文仍有残余。这使 1.2 kbps 成为原创系统有意义的内容保真下界,而不是只以听起来像无线电为成功。
- 本机 encode/decode 均显著快于实时,但属于 macOS arm64、短句、冷进程条件。D4 已证明 Linux/macOS 可互解,却也发现 2400/3200 bitstream 与所有 decoder PCM 并非跨 build bit-exact;部署回归应比较合同和容差,不能无条件锁死 PCM hash。
明确没有证明什么¶
D5 没有复现历史 pitch/LSP/听测结果,也没有验证噪声、丢包、误码、FEC、无线信道、对话、长音频、关键实体或六语泛化。 ViSQOL/STOI/ASR/speaker/F0 等自动指标不等于人的内容、身份、表达和总体偏好;中文/英文语料之间也不是受控语言因果实验。个人录音仍是可选 D6 扩展,不阻塞本 Gate。 D5 passed 的准确含义只是:Codec 2 的公共 180-row benchmark 档案完整、来源明确,并能由已冻结父证据确定性重建。
Codec 2 D5 · 公共统一横评结果¶
状态:passed。这是已完成共同实验的确定性模型级归档,不是一次新的 codec、ASR 或质量指标运行。 180 个逐样本对象与父证据逐字段相等;三档 aggregate 与固定 bootstrap CI 原样保留。
三档全体结果(30 英 + 30 中)¶
| 点位 | representation bps | serialized payload bps | EVSC2V01 overhead | ViSQOL | STOI | Whisper ΔER en/zh | SenseVoice ΔER en/zh | WavLM | Resemblyzer | enc ×RT | dec ×RT |
|---|---|---|---|---|---|---|---|---|---|---|---|
codec2_1200bps |
1197.5 | 1733.3 | +44.74% | 3.7105 | 0.6302 | +5.06%/+6.42% | +3.03%/+5.00% | 0.9078 | 0.7111 | 17.66 | 21.93 |
codec2_2400bps |
2398.4 | 2934.8 | +22.37% | 3.7338 | 0.6317 | +3.18%/+2.48% | +1.09%/+2.86% | 0.9134 | 0.7181 | 18.35 | 22.29 |
codec2_3200bps |
3197.9 | 3734.4 | +16.78% | 3.7400 | 0.6416 | +2.27%/+0.07% | +0.61%/+1.03% | 0.9322 | 0.7470 | 18.00 | 23.34 |
ΔER = decoded error rate − reference-audio error rate;正值表示 codec 输出让该 ASR 更难识别。绝对数必须与 reference-audio baseline 一起解释。
语言条件与 95% CI¶
| 点位 | 语言 | ViSQOL mean [95% CI] | STOI mean [95% CI] | Whisper ΔER [95% CI] | SenseVoice ΔER [95% CI] |
|---|---|---|---|---|---|
codec2_1200bps |
en | 3.2668 [3.1478, 3.3819] | 0.6457 [0.6330, 0.6591] | +0.0506 [+0.0149, +0.0908] | +0.0303 [+0.0103, +0.0520] |
codec2_1200bps |
zh | 4.1542 [4.0148, 4.2934] | 0.6147 [0.6028, 0.6272] | +0.0642 [-0.0032, +0.1291] | +0.0500 [+0.0173, +0.0897] |
codec2_2400bps |
en | 3.2985 [3.1813, 3.4134] | 0.6470 [0.6340, 0.6613] | +0.0318 [+0.0051, +0.0648] | +0.0109 [-0.0097, +0.0397] |
codec2_2400bps |
zh | 4.1690 [4.0283, 4.3063] | 0.6164 [0.6034, 0.6300] | +0.0248 [-0.0332, +0.0808] | +0.0286 [+0.0036, +0.0588] |
codec2_3200bps |
en | 3.3274 [3.2094, 3.4381] | 0.6580 [0.6457, 0.6707] | +0.0227 [+0.0015, +0.0488] | +0.0061 [-0.0066, +0.0220] |
codec2_3200bps |
zh | 4.1526 [4.0201, 4.2819] | 0.6252 [0.6116, 0.6390] | +0.0007 [-0.0583, +0.0557] | +0.0103 [-0.0091, +0.0323] |
四本不能混的账¶
- raw representation 平均约 1.198/2.398/3.198 kbps;这是 mode frame bits 除以真实时长;
- 完整 EVSC2V01 payload 平均约 1.733/2.935/3.734 kbps,短句固定元数据使相对 overhead 为 +44.74%/+22.37%/+16.78%;
- common benchmark 的 endpoint model state 为 0,只表示没有外部 checkpoint;
- D2 仍确认所选路径含 10,752 个 data-derived VQ floats(43,008 bytes)、452 个 designed scalar-grid 元素,程序/动态库 footprint 另算。Codec 2 不能被描述为“零参数”或“零模型成本”。
读数与研究启发¶
- 1.2→2.4 kbps 的全体 ViSQOL 只从 3.7105 到 3.7338、STOI 从 0.6302 到 0.6317;3.2 kbps 才把 STOI、WavLM、Resemblyzer 和双 ASR ΔER 同时推向更好方向。这说明 Codec 2 的低档瓶颈不只是再多几个 bits,而是 8-kHz 窄带、参数表示与 mode-specific quantiser/synthesis 的共同限制。
- SI-SNR 在三档仍约为 −33 到 −30 dB,而 ViSQOL 约 3.71–3.74。参数语音会重建相位和谐波细节,不追求 sample waveform 对齐;因此未来 AI 进化若把 SI-SNR 当唯一 reward,会系统性惩罚可懂但波形不同的好方案。
- 中文 ViSQOL 约 4.15、英文约 3.27,差距在三档都存在;它更可能反映两个语料分布、录音条件和 evaluator 行为,而不是‘Codec 2 天生更适合中文’。报告必须保留分语言表,不能只看 60 句平均。
- 1.2 kbps 下 Whisper/SenseVoice 的 en/zh ΔER 都明显为正;提高到 3.2 kbps 后中文增量接近零、英文仍有残余。这使 1.2 kbps 成为原创系统有意义的内容保真下界,而不是只以听起来像无线电为成功。
- 本机 encode/decode 均显著快于实时,但属于 macOS arm64、短句、冷进程条件。D4 已证明 Linux/macOS 可互解,却也发现 2400/3200 bitstream 与所有 decoder PCM 并非跨 build bit-exact;部署回归应比较合同和容差,不能无条件锁死 PCM hash。
明确没有证明什么¶
D5 没有复现历史 pitch/LSP/听测结果,也没有验证噪声、丢包、误码、FEC、无线信道、对话、长音频、关键实体或六语泛化。 ViSQOL/STOI/ASR/speaker/F0 等自动指标不等于人的内容、身份、表达和总体偏好;中文/英文语料之间也不是受控语言因果实验。个人录音仍是可选 D6 扩展,不阻塞本 Gate。 D5 passed 的准确含义只是:Codec 2 的公共 180-row benchmark 档案完整、来源明确,并能由已冻结父证据确定性重建。