docs(tools): close the NEL half of the raw pass-through and reflow
Four non-blocking review suggestions, all prose plus one assertion. `UNPRINTABLE`'s new sentence named three characters but only two raw-reach points, leaving "and NEL?" open; it now says all three reach text through `pyScalar`, and how the description path handles each. `pyScalar`'s raw-pass-through list already covered NEL under "the C1 controls", and the test now pins it alongside LS and PS, so the docstring's claim has a mechanical check for every character it names. The test title said "paragraph separators" for a pair whose first member is LINE SEPARATOR. Two docstring paragraphs are reflowed to the file's ~80 columns after the earlier inserts left short lines. The note's CPython-floor obligation gains a second axis: the `typing` names the block spells (`TypedDict` 3.8, `NotRequired` 3.11, `A | B` annotations 3.10) are definition-time evaluation floors, not parse floors, so the floor PR does not read "parseable on the supported range" as "executable on it".
This commit is contained in:
5 files changed
+31
-27
No files matched your search
@@ -2,5 +2,5 @@
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write .agents/notes/implemented/feature/2026-07-31-code-mode-language-dispatch.md
|
||||
2026-07-31-code-mode-language-dispatch.md: c46f64b704daa5d6cededb6be96f64e825e59a5d
|
||||
2026-07-31-code-mode-language-dispatch.zh.md: 1851cc18780c1cdf624cb0285669eb0b7b53f14a
|
||||
2026-07-31-code-mode-language-dispatch.md: e7adc5386e101bd02aba525a22070f5cac3d840f
|
||||
2026-07-31-code-mode-language-dispatch.zh.md: ebfb228aaca7a3aa2a3a7b9f44977f1b0cebe045
|
||||
@@ -43,4 +43,4 @@ The cost is that the Python branch of both tables is unreachable on this base: `
|
||||
|
||||
Two runtime contracts the Python SDK text asserts are owed by that same backend PR. First, the instructions tell the model that exactly `tools` and `ToolCallError` are bound and that the declared `TypedDict` classes are not, so the backend must inject those two names — with `ToolCallError.toolName` populated per the seam's `errorClass` contract — and must NOT bind the declared class names into the program's globals; injecting them "helpfully" would make the SDK text false. Second, the language has to be bound to the request: `requireCodeRuntime` resolves `ctx.codeRuntime` separately at assembly and at `run_code` execution, so a reload that swapped the runtime between those two points would hand a program written against one flavor to the other. The split is finer than those two points — `run_code`'s `description` and `parameters` getters each call `resolveFlavor(peekRuntime())`, and `schemaOf` destructures both, so one projection reads the runtime twice; both reads are for `run_code`'s own schema, since the getters are installed on that one definition and every other definition carries plain data properties. A reload between those two reads yields a single schema whose two halves name different languages. Neither is reachable here — one published backend means both reads return the same flavor and no program ever runs against this renderer's output — and the cross-language rejection is not testable until a second language exists.
|
||||
|
||||
Third, that PR owns the CPython floor, and with it the renderer's Unicode-table skew. Four expressions read the running engine's tables (Node 22.23.1: Unicode 17.0) while the interpreter uses its own (CPython 3.9.6: 13.0.0): `isBareIdentifier`'s `IDENTIFIER`, and `camelCase`'s split set, head test, and `toUpperCase()`. An interpreter older than the engine is the failing direction — the engine emits a character its tokenizer refuses, taking the whole block down — and it arrives by three independent paths. Through the predicate, a bare method or field name carrying a character added between the two versions — to `XID_Start` at its head, or to `XID_Continue` in any tail position, the middle of a name included. Through `camelCase`'s XID reads, a class name, which reaches emitted text whenever any object shape in the tool's schema declares a `TypedDict`, and which the predicate's verdict on the tool name does not gate: `zz-` plus U+1E4D0 never reaches the predicate's skew, since the `-` rejects it outright, yet it still declares `class Zz𞓐xArgs`. Through the case mapping, a class name derived from a tool the predicate accepted — a different table and a wider window than XID membership: U+019B is XID_Start and NFKC-stable, so `async def ƛ` compiles on 3.9.6, but Node uppercases it to U+A7DC (unassigned there; CPython's own `.upper()` is the identity) and `class Args` fails with `invalid non-printable character U+A7DC`. The exposure window is the characters and mappings that changed between the two versions, so the PR that names a supported CPython range must decide explicitly between accepting it and pinning all four read points to tables for that floor — pinning the predicate alone leaves both class-name paths open. Nothing here can decide it: the floor does not exist yet, and a table pinned to a guess would be a deployment-varying constant with no configurability behind it.
|
||||
Third, that PR owns the CPython floor, and with it the renderer's Unicode-table skew. Four expressions read the running engine's tables (Node 22.23.1: Unicode 17.0) while the interpreter uses its own (CPython 3.9.6: 13.0.0): `isBareIdentifier`'s `IDENTIFIER`, and `camelCase`'s split set, head test, and `toUpperCase()`. An interpreter older than the engine is the failing direction — the engine emits a character its tokenizer refuses, taking the whole block down — and it arrives by three independent paths. Through the predicate, a bare method or field name carrying a character added between the two versions — to `XID_Start` at its head, or to `XID_Continue` in any tail position, the middle of a name included. Through `camelCase`'s XID reads, a class name, which reaches emitted text whenever any object shape in the tool's schema declares a `TypedDict`, and which the predicate's verdict on the tool name does not gate: `zz-` plus U+1E4D0 never reaches the predicate's skew, since the `-` rejects it outright, yet it still declares `class Zz𞓐xArgs`. Through the case mapping, a class name derived from a tool the predicate accepted — a different table and a wider window than XID membership: U+019B is XID_Start and NFKC-stable, so `async def ƛ` compiles on 3.9.6, but Node uppercases it to U+A7DC (unassigned there; CPython's own `.upper()` is the identity) and `class Args` fails with `invalid non-printable character U+A7DC`. The exposure window is the characters and mappings that changed between the two versions, so the PR that names a supported CPython range must decide explicitly between accepting it and pinning all four read points to tables for that floor — pinning the predicate alone leaves both class-name paths open. Nothing here can decide it: the floor does not exist yet, and a table pinned to a guess would be a deployment-varying constant with no configurability behind it. A second axis rides along with the floor and is not one of the four: the `typing` names the block spells. `TypedDict` needs 3.8, `NotRequired` 3.11, and a `A | B` annotation evaluates only on 3.10. These are not parse failures — the block parses on any version, which is the standard the `MAX_LIST_NESTING` cap serves — but definition-time evaluation failures, and nothing in the product evaluates this text. Recording them with the read points keeps "parseable on the supported range" from being read as "executable on it".
|
||||
@@ -43,4 +43,4 @@ Code Mode 只生成一种 SDK 形态:TypeScript。`ToolRegistry` 为 `tools:sd
|
||||
|
||||
Python SDK 文本断言的两条运行时契约同样归属那个 backend PR。其一,说明文字告诉模型运行时恰好绑定 `tools` 与 `ToolCallError` 两个名字、所声明的 `TypedDict` 类不绑定,因此后端必须注入这两个名字(并按 seam 的 `errorClass` 契约填充 `ToolCallError.toolName`),且**不得**把所声明的类名绑进程序全局——「好心」注入会使这段 SDK 文本变成假话。其二,语言必须绑定到请求上:`requireCodeRuntime` 在组装时与 `run_code` 执行时分别解析 `ctx.codeRuntime`,若在这两点之间发生重载并换掉运行时,就会把针对一种形态写成的程序交给另一种形态执行。分裂比这两点更细——`run_code` 的 `description` 与 `parameters` 两个 getter 各自调用 `resolveFlavor(peekRuntime())`,而 `schemaOf` 会解构这两个字段,因此一次投影读两次运行时;两次都属于 `run_code` 自己的 schema,因为这两个 getter 只装在那一个 definition 上,其余 definition 携带的都是普通数据属性。在这两次读取之间重载会产出单个 schema 的两半分属不同语言。两者在此处都不可达——只有一个已发布后端意味着两次读取返回同一形态,且没有任何程序会针对本渲染器的输出运行——而跨语言拒绝在第二门语言存在之前也无法测试。
|
||||
|
||||
其三,那个 PR 拥有 CPython 版本下限,连带拥有本渲染器的 Unicode 表偏斜。有四处表达式读所运行引擎的表(Node 22.23.1:Unicode 17.0),而解释器用它自己的表(CPython 3.9.6:13.0.0):`isBareIdentifier` 的 `IDENTIFIER`,以及 `camelCase` 的切分集、头部测试与 `toUpperCase()`。解释器旧于引擎是会失败的那个方向——引擎发出的字符被其 tokenizer 拒收,整个块随之不可解析——而它经三条独立路径抵达。经判据抵达的是裸发的方法名或字段名,其中带有一个在两个版本之间新增的字符——首位加进 `XID_Start`,或尾部任意位置(含名字中部)加进 `XID_Continue`。经 `camelCase` 的 XID 读取抵达的是类名:只要工具 schema 中有任一对象形态声明 `TypedDict`,该类名就进入发出的文本,且判据对工具名的裁决并不对它设闸——工具名 `zz-` 加 U+1E4D0 因 `-` 被判据直接拒绝、从不触及那里的偏斜,却照样声明 `class Zz𞓐xArgs`。经大写映射抵达的是由判据已接受的工具派生出的类名——这是另一张表,窗口也比 XID 归属更宽:U+019B 既是 XID_Start 又 NFKC 稳定,故 `async def ƛ` 在 3.9.6 上可编译,但 Node 将其大写为 U+A7DC(在那里未分配;CPython 自己的 `.upper()` 在此是恒等),于是 `class Args` 以 `invalid non-printable character U+A7DC` 失败。暴露窗口是两个版本之间发生变化的那些字符与映射,所以宣布支持某个 CPython 范围的那个 PR 必须在「接受该暴露」与「按该下限的表钉住全部四个读取点」之间显式作出决定——只钉判据会同时留下两条类名路径。此处无法决定:下限尚不存在,而按猜测钉死一张表会成为一个随部署而变、却没有可配置性支撑的常量。
|
||||
其三,那个 PR 拥有 CPython 版本下限,连带拥有本渲染器的 Unicode 表偏斜。有四处表达式读所运行引擎的表(Node 22.23.1:Unicode 17.0),而解释器用它自己的表(CPython 3.9.6:13.0.0):`isBareIdentifier` 的 `IDENTIFIER`,以及 `camelCase` 的切分集、头部测试与 `toUpperCase()`。解释器旧于引擎是会失败的那个方向——引擎发出的字符被其 tokenizer 拒收,整个块随之不可解析——而它经三条独立路径抵达。经判据抵达的是裸发的方法名或字段名,其中带有一个在两个版本之间新增的字符——首位加进 `XID_Start`,或尾部任意位置(含名字中部)加进 `XID_Continue`。经 `camelCase` 的 XID 读取抵达的是类名:只要工具 schema 中有任一对象形态声明 `TypedDict`,该类名就进入发出的文本,且判据对工具名的裁决并不对它设闸——工具名 `zz-` 加 U+1E4D0 因 `-` 被判据直接拒绝、从不触及那里的偏斜,却照样声明 `class Zz𞓐xArgs`。经大写映射抵达的是由判据已接受的工具派生出的类名——这是另一张表,窗口也比 XID 归属更宽:U+019B 既是 XID_Start 又 NFKC 稳定,故 `async def ƛ` 在 3.9.6 上可编译,但 Node 将其大写为 U+A7DC(在那里未分配;CPython 自己的 `.upper()` 在此是恒等),于是 `class Args` 以 `invalid non-printable character U+A7DC` 失败。暴露窗口是两个版本之间发生变化的那些字符与映射,所以宣布支持某个 CPython 范围的那个 PR 必须在「接受该暴露」与「按该下限的表钉住全部四个读取点」之间显式作出决定——只钉判据会同时留下两条类名路径。此处无法决定:下限尚不存在,而按猜测钉死一张表会成为一个随部署而变、却没有可配置性支撑的常量。还有第二条轴随该下限一同确定,且不属于那四个读取点:本块所拼写的 `typing` 名字。`TypedDict` 需要 3.8,`NotRequired` 需要 3.11,而 `A | B` 形式的注解只在 3.10 及以上才可求值。这些不是解析失败——本块在任何版本上都能解析,这正是 `MAX_LIST_NESTING` 上限所服务的标准——而是定义期求值失败,且产品中没有任何东西会求值这段文本。把它们与那四个读取点记在一起,可避免把「在所支持范围上可解析」读成「在其上可执行」。
|
||||
@@ -172,7 +172,9 @@ interface RenderState {
|
||||
* string at run time but do not end a physical line in source — measured on
|
||||
* CPython 3.9.6 and 3.12.13, each accepted in both positions with the value
|
||||
* round-tripping — so they are safe raw wherever they reach emitted text
|
||||
* unescaped, which for LS and PS is {@link pyScalar}'s `JSON.stringify`.
|
||||
* unescaped, which for all three is {@link pyScalar}'s `JSON.stringify`: the
|
||||
* `description` path escapes NEL under the class above and folds LS and PS in
|
||||
* {@link describe}'s `\s+` collapse, both of them being ECMAScript `\s`.
|
||||
*/
|
||||
const UNPRINTABLE = /[\u0000-\u0008\u000e-\u001f\u007f-\u009f]/g
|
||||
|
||||
@@ -404,24 +406,23 @@ function childClassName(base: string, segment: string): string {
|
||||
* code point CPython refuses anywhere in source — NUL among the C0 controls,
|
||||
* and the whole D800–DFFF unpaired-surrogate block, escaped under ES2019
|
||||
* well-formed stringification, which the engines range guarantees — and the
|
||||
* ones that break this line in particular,
|
||||
* a bare `"` closing the literal early, a trailing odd backslash eating the
|
||||
* closing quote, and a bare LF/CR ending it before its terminator. The
|
||||
* `description` path carries {@link UNPRINTABLE} and {@link LONE_SURROGATE}
|
||||
* because nothing quotes it, and folds newlines in {@link describe}.
|
||||
* ones that break this line in particular, a bare `"` closing the literal
|
||||
* early, a trailing odd backslash eating the closing quote, and a bare LF/CR
|
||||
* ending it before its terminator. The `description` path carries
|
||||
* {@link UNPRINTABLE} and {@link LONE_SURROGATE} because nothing quotes it,
|
||||
* and folds newlines in {@link describe}.
|
||||
*
|
||||
* That leans on a coincidence worth naming: every escape `JSON.stringify` can
|
||||
* emit (`\"`, `\\`, `\b`, `\f`, `\n`, `\r`, `\t`, `\uXXXX`) is also a Python
|
||||
* escape denoting the same character, so the emitted `Literal[...]` both
|
||||
* parses and decodes back to the value the schema declared. DEL, the C1
|
||||
* controls, and LS/PS (U+2028/U+2029) do reach it raw — legal but invisible,
|
||||
* byte-for-byte as in the TS flavor; escaping them is a both-flavors change.
|
||||
* LS and PS are legal here for the reason {@link UNPRINTABLE} records: they
|
||||
* are `str.splitlines()` boundaries, not tokenizer line terminators. The
|
||||
* subscript tool-name
|
||||
* comment quotes its name through its own call to the same `JSON.stringify`,
|
||||
* never through this function, and inherits both halves — escapes and
|
||||
* pass-throughs alike.
|
||||
* controls (NEL among them), and LS/PS (U+2028/U+2029) do reach it raw —
|
||||
* legal but invisible, byte-for-byte as in the TS flavor; escaping them is a
|
||||
* both-flavors change. Those last three are legal here for the reason
|
||||
* {@link UNPRINTABLE} records: they are `str.splitlines()` boundaries, not
|
||||
* tokenizer line terminators. The subscript tool-name comment quotes its name
|
||||
* through its own call to the same `JSON.stringify`, never through this
|
||||
* function, and inherits both halves — escapes and pass-throughs alike.
|
||||
*/
|
||||
function pyScalar(value: JsonSchemaScalar): string {
|
||||
if (value === true) return 'True'
|
||||
|
||||
@@ -67,17 +67,20 @@ describe('jsonSchemaToPy', () => {
|
||||
expect(jsonSchemaToPy({ type: 'string', const: 'ends\\' })).toBe(String.raw`Literal["ends\\"]`)
|
||||
})
|
||||
|
||||
it('passes the paragraph separators through raw, which CPython does not treat as line terminators', () => {
|
||||
// `JSON.stringify` escapes LF and CR but not LS/PS (U+2028/U+2029), which
|
||||
// is safe here and not by accident: they are `str.splitlines()` boundaries,
|
||||
// not tokenizer line terminators, so they end neither a string literal nor
|
||||
// a `#` comment — measured on CPython 3.9.6 and 3.12.13. Pinning the raw
|
||||
// form keeps a later "escape them for symmetry with LF" change from
|
||||
// landing as a silent both-flavors divergence from `ts-types`.
|
||||
// Escapes below — the two forms denote the same bytes, and neither
|
||||
// character has a visible width.
|
||||
it('passes the line and paragraph separators through raw, which CPython does not treat as line terminators', () => {
|
||||
// `JSON.stringify` escapes LF and CR but not NEL (U+0085), LS (U+2028), or
|
||||
// PS (U+2029), which is safe here and not by accident: those three are
|
||||
// `str.splitlines()` boundaries, not tokenizer line terminators, so they
|
||||
// end neither a string literal nor a `#` comment — measured on CPython
|
||||
// 3.9.6 and 3.12.13. Pinning the raw form keeps a later "escape them for
|
||||
// symmetry with LF" change from landing as a silent both-flavors
|
||||
// divergence from `ts-types`. Escapes below — the two forms denote the
|
||||
// same bytes, and none of the three has a visible width.
|
||||
expect(jsonSchemaToPy({ type: 'string', const: 'a\u2028b' })).toBe('Literal["a\u2028b"]')
|
||||
expect(jsonSchemaToPy({ type: 'string', enum: ['a\u2029b'] })).toBe('Literal["a\u2029b"]')
|
||||
// NEL is inside `UNPRINTABLE`'s class, so the description path escapes it;
|
||||
// this is the one route that carries it raw.
|
||||
expect(jsonSchemaToPy({ type: 'string', const: 'a\u0085b' })).toBe('Literal["a\u0085b"]')
|
||||
})
|
||||
|
||||
it('emits exact digits for a beyond-safe-range integer literal', () => {
|
||||
|
||||
Reference in New Issue
Block a user