Most of what ships here exists because a number I trusted turned out to be wrong, and I wanted the next one to fail loudly instead of quietly.
|
Verify the instrument before you charge the failure to the system under test. A measurement-validity checklist for data/ML work. Every entry is a real failure I first booked to the wrong cause. |
Turn a lesson learned into a guardrail that cannot silently rot. Semgrep rules, Soda checks, and lint gates locked by offline known-answer selftests — the arbiter can't drift without CI noticing. |
|
A backtester that refuses to run an experiment with no random-control baseline, no out-of-sample split, and no real two-sided costs. It falsified six strategy families, including the ones I wanted to be true. |
Measures what's wrong with a résumé instead of opining about it — layout metrics, line-break calibration, copy lint. Every rule carries the measurement that justifies it. |
Also here: investment_research · polypost · windup-asset-lab
188 merged pull requests across a production menu-OCR and data pipeline covering ten markets — extraction, brand entity matching, warehouse write paths, and cost. The three I'd point at:
| What it caught | |
|---|---|
| A CI gate | Asserts a promote step actually succeeded. Five consecutive failed runs had reported green. |
| A quality gate | Found after a 16-hour run produced a 0% pair rate that nobody noticed for a month. Deliberately fail-open — halting a pipeline on a quality warning is a worse failure than the one it catches. |
| A transport gate | The batch API returned 404 on the pinned model while every read-only check — model list, capability flags, batch list — kept returning 200. Only the create call refused. |
I care about the denominator. Most coverage arguments I've been in were settled by finding out what was actually in the denominator, not by improving the model.
Graduating June 2027. Looking for data / AI engineering roles — Shanghai, Hangzhou, or remote-friendly. Comfortable working entirely in English.
我做的大多是第二种。这里的东西基本都源于同一件事:某个我信过的数字后来被证明是错的, 于是我想让下一个错数响一点地失败,而不是安静地混过去。
|
先验仪器,再给被测系统记账。数据 / ML 工程的测量效度清单。 每一条都是我先归错因、后来才查出真因的真实事故。 |
把一次教训变成不会静默腐烂的闸口。semgrep 规则、Soda 质检、三个 lint 闸, 判定器被离线已知答案自测锁住 —— 判官自己漂移了 CI 会先叫。 |
|
一个拒绝运行的回测框架:没声明随机对照基线、样本外切分和真实双边成本,就报错中止。 它证伪了六类策略,包括我希望是真的那几类。 |
量得出简历哪里有问题,而不只是给意见 —— 版面度量、折行标定、文案 lint。 每条规则都带着支撑它的那次实测。 |
在一条覆盖十个市场的生产菜单 OCR 与数据管道上合并了 188 个 PR —— 抽取、品牌实体匹配、数仓写入路径、成本。 最值得一提的三个:
| 它拦住了什么 | |
|---|---|
| 一个 CI 闸 | 断言 promote 步骤真的成功了。在此之前,连续五次失败的运行全部报绿。 |
| 一个质量闸 | 起因是一次 16 小时的运行产出 0% 配对率,一个月没人发现。它故意 fail-open —— 质量告警去中断管线,会造出一个比原问题更糟的失败模式。 |
| 一个传输资格闸 | batch 接口在钉住的模型上返回 404,而所有只读检查——模型列表、能力标记、batch 列表——全部返回 200,只有 create 拒绝。 |
我在意分母。 我经历过的覆盖率争论,绝大多数最后是靠搞清楚分母里到底装了什么解决的,不是靠把模型调好。
2027 年 6 月毕业,找数据 / AI 工程方向的岗位 —— 上海、杭州,或可远程。全英文工作环境没问题。



