Implementation details, platform quirks, and edge cases encountered during development.
Claude Code passes hook input as JSON on stdin. On Windows, two issues compound:
-
uv run(without--script) consumes PostToolUse stdin. The project sync phase closes or drains the stdin pipe before the Python script starts. PreToolUse works because uv's first-run environment setup completes before stdin is read. Workaround:uv run --scriptskips project discovery entirely. -
Python defaults to GBK codepage for stdin. Claude Code sends UTF-8 JSON, but
sys.stdin.read()decodes via the system codepage (GBK on Chinese Windows). Characters outside GBK range causeUnicodeDecodeError. Workaround:sys.stdin.buffer.read().decode("utf-8")— binary read then explicit UTF-8 decode. -
Stdin must be read before heavy imports. Some imports or uv initialization may interfere with the pipe. The script reads
sys.stdin.buffer.read()on the very first line afterimport sys.
On Windows, python3 is typically a Microsoft Store stub (exit code 49, opens Store). uv run --script bypasses this entirely — it manages its own Python discovery.
Claude Code's Edit tool may convert CRLF to LF (#38887). The hook detects the dominant line ending style before conversion and restores it after. Implementation: all line endings are unified to \n first, then expanded to \r\n if the original was CRLF. This correctly handles mixed-EOL files.
Windows antivirus or indexing services may temporarily lock files. The hook handles this in two places:
atomic_write(): ifos.replace()fails, the temp file is cleaned up and the original file is preserved.delete_cache(): returnsFalseon failure; the Post handler logs a warning and the Pre handler will detect the stale cache on next invocation.
chardet 7.x is a complete rewrite (Mar 2026) with a cosine-similarity bigram-model scoring system. Its confidence values are not directly comparable to chardet 5.x: same content gives very different numbers, and the calibration varies by encoding (driven by bigram inventory size in each model).
| Encoding | chardet 5.2.0 confidence | chardet 7.4.3 confidence (typical, 5KB sample) |
|---|---|---|
| GBK / GB2312 | 0.99 | 0.59–0.67 |
| Big5 | 0.99 | 0.37–0.40 (large bigram inventory geometry, see chardet design doc) |
| GB18030 | 0.99 | 0.14–0.25 |
| EUC-JP | 0.99 | 0.83–0.91 |
| Shift_JIS | 0.99 | 0.83 (returned as cp932, not SHIFT_JIS) |
| EUC-KR | 0.99 | 0.85 (returned as CP949, not EUC-KR) |
Other behavioral changes in 7.x that block a drop-in upgrade:
- Built-in binary detection at pipeline stage 5 returns
encoding=Nonefor binaries (could replace binaryornot) - Default encoding names changed for some codecs —
shift_jis→cp932,euc-kr→CP949(deliberate for the latter; arguably a missing_COMPAT_NAMESmapping for the former) max_bytes,encoding_era,include_encodings,prefer_superset,compat_names— new tuning parameters
Pinned to >=5,<6 via PEP 723 inline metadata until a deliberate 7.x migration redesigns the encoding-set, threshold, and binary-check layers together. The uv run --script creates an isolated environment — the pinned version does not affect the user's project dependencies.
Tested with real .lib and .dll files:
| File | chardet result | Would trigger hook? |
|---|---|---|
iconv.lib (3KB) |
Windows-1252, 0.73 | Yes — dangerous |
pthreadVC2.dll (55KB) |
KOI8-R, 0.61 | Yes — dangerous |
gwschedule_32_d.lib (4KB) |
Windows-1252, 0.73 | Yes — dangerous |
dscompress.dll (119KB) |
None, 0.00 | No |
Windows-1252 is a single-byte encoding that can decode almost any byte sequence without error. Without binary detection, the hook would "successfully" convert binary files, and Claude's subsequent edit would corrupt them irreversibly.
binaryornot uses a trained decision tree with 24 features including CJK encoding validity checks, Shannon entropy, magic signatures, and null byte ratios. It correctly identifies all tested binary files but false-positives in two patterns:
- Short CJK files (<500 bytes Shift_JIS / EUC-JP / EUC-KR / Big5 / GB18030, observed at 50–240 bytes in PoC).
- Mixed Cyrillic + ASCII files — e.g., source code with CP1251 comments. Empirically a 115-byte C source with a single CP1251 comment line is flagged binary even though
chardetreturnswindows-1251at 0.99 confidence.
To resolve this without losing binaryornot's protection on Windows-1252 binaries, the Pre hook uses a structural-trust short-circuit:
encoding, confidence = chardet.detect(...)
if _strip_enc(norm) not in _RESTORE_SET: # outside our supported set: skip
return
if confidence < 0.5: # confidence floor: skip
return
if _strip_enc(norm) not in STRUCTURAL_TRUSTED: # only run binaryornot if not trusted
if is_binary(path):
returnSTRUCTURAL_TRUSTED = {gbk, gb18030, big5, big5hkscs, shiftjis, eucjp, euckr, iso2022jp, windows1251}. These encodings have strict structural rules (multi-byte lead/trail ranges for CJK; characteristic Cyrillic byte distribution for cp1251) that chardet's probers verify directly — a real binary cannot satisfy them across hundreds of bytes at confidence ≥ 0.5. binaryornot remains the binary check for Windows-1252 / ISO-8859-1, where chardet alone cannot rule out binary (single-byte Latin encodings can decode any bytes).
Validated on 100+ fixtures (our PoC + binaryornot's own test set + Cyrillic 8-size matrix in _backup/fixtures/cyrillic_sizes/): all supported text files correctly converted with byte-identical round-trip, all binaries correctly rejected, zero corruption.
chardet reports GB2312 but Python's gb2312 codec is stricter than gbk. Real-world files labeled GB2312 often contain GBK-range characters. Mapping to the superset (gbk) avoids UnicodeEncodeError during restore.
Similarly, ISO-8859-1 → Windows-1252 follows the HTML specification's behavior. The 0x80-0x9F range in real files almost always contains CP1252 characters (€, curly quotes, em dash), not C1 control characters.
normalize_encoding returns chardet's name verbatim for the non-aliased path:
def normalize_encoding(enc: str) -> str:
stripped = _strip_enc(enc)
for orig, alias in ENCODING_ALIASES.items():
if stripped == _strip_enc(orig):
return alias
return enc # NOT _strip_enc(enc)Returning the dash-stripped form ("windows1251", "windows1252", "iso88591") breaks codec lookup. Python's codec normalizer replaces - with _ but cannot recover a missing separator: windows-1251 / windows_1251 / cp1251 all resolve to the cp1251 codec, but windows1251 raises LookupError.
The aliased path (gb2312 → gbk, iso-8859-1 → windows-1252) was unaffected because the alias value is returned verbatim with its dashes intact. Direct chardet hits on windows-1252 silently failed in convert_file until #3 added windows-1251 to RESTORE_ENCODINGS and exposed the same pattern across every Cyrillic test size — surfacing the latent bug for both encodings simultaneously.
Without session isolation:
Session A: Read → convert GBK→UTF-8 → cache
Session B: Read → sees cache exists → skip (uses A's cache)
Session A: Edit → Post restore → delete cache
Session B: Edit → Post → no cache → cannot restore!
With session isolation (<tmpdir>/.cc_encoding_cache/<session_id>/), each session manages its own cache independently.
Pre hook validates cache on every access:
- Cache exists for this file in this session?
- Can the file be decoded as UTF-8? (
raw.decode("utf-8"))- Yes → file is in converted state, cache is valid, skip
- No → file was already restored (stale cache from failed delete), remove cache, re-convert
Edge case: Windows-1252 files whose original bytes are valid UTF-8 (e.g., C2 A9 = © in UTF-8, © in Windows-1252). The stale cache won't self-heal because the file passes the UTF-8 decode check. This requires both delete_cache failure AND the original bytes to be coincidentally valid UTF-8 — a narrow edge case that doesn't affect CJK encodings.
- Session directories older than 24 hours are removed on each Pre hook invocation
- Current session is always skipped during cleanup (compared by sanitized session_id)
- Empty session directories are removed after the last cache file is deleted
The initial design converted files in PreToolUse for Edit only:
- Pre(Edit): detect GBK → convert to UTF-8 → cache
- Claude edits
- Post(Edit): UTF-8 → GBK
This failed because Claude Code v2.1.90+ silently accepts file modifications from hooks without re-reading content. Claude's Edit uses its in-memory content from the last Read (which was garbled GBK-as-UTF-8), writes it back, and the file gets corrupted with U+FFFD.
The permissionDecision: "deny" mechanism was tried to force Claude to re-Read after conversion. This failed because:
- Claude retries the Edit without actually re-Reading (uses stale in-memory content)
- PostToolUse fires for denied edits too, prematurely consuming the cache
- The file bounces between GBK and UTF-8 in a loop
Converting at Read time ensures Claude's first in-memory view is already correct UTF-8. The Pre hook matcher includes Read|Edit|Write for Pre (convert before any file access) and Edit|Write for Post (restore after modification).
Post only fires after Edit/Write. If Claude reads a file but never edits it in that turn, the file is left in UTF-8 form and the cache entry is never consumed. To recover, a Stop hook runs encoding_guard.py restore-all at the end of every Claude turn: it enumerates the session cache directory (Stop stdin has no tool_input.file_path) and restores each still-cached file through the same convert_file helper that handle_post uses. Cache validation goes through load_cache so the Stop path applies the same _RESTORE_SET / line_ending / type checks as Post. (Approach contributed by @lbresler in #2.)
Accepted trade-off — multi-session race. When two CC instances operate on the same file, the restore triggered by one session's Stop can overwrite a pending Edit in another session. Session-isolated caches prevent most cross-contamination, but a narrow race window remains. Concurrent CC use is rare enough that this trade-off is accepted in exchange for Read-without-Edit recovery in the common single-session case.
All file writes use atomic_write():
- Create temp file in the same directory (
tempfile.mkstemp) - Write data,
flush(),os.fsync() os.replace()atomically replaces the original
If step 2 or 3 fails, the temp file is cleaned up and the original file is preserved. This prevents data loss from disk-full errors, process interruption, or permission issues.
Empirically verified against Claude Code 2.1.132 via an instrumented log-only hook in a separate playground project. Documented here as a stable reference for future hook design (read-only redirect, raw-byte backup, ACL gate, etc.).
PreToolUse for Read:
{
"session_id": "...",
"transcript_path": "C:\\Users\\...\\<session_id>.jsonl",
"cwd": "...",
"permission_mode": "default",
"hook_event_name": "PreToolUse",
"tool_name": "Read",
"tool_input": {"file_path": "..."},
"tool_use_id": "toolu_..."
}PostToolUse for Read:
{
"session_id": "...",
"transcript_path": "...",
"cwd": "...",
"permission_mode": "default",
"hook_event_name": "PostToolUse",
"tool_name": "Read",
"tool_input": {"file_path": "..."},
"tool_response": {
"type": "text",
"file": {
"filePath": "...",
"content": "...",
"numLines": 2,
"startLine": 1,
"totalLines": 2
}
},
"tool_use_id": "toolu_...",
"duration_ms": 6
}Stop:
{
"session_id": "...",
"transcript_path": "...",
"cwd": "...",
"permission_mode": "default",
"hook_event_name": "Stop",
"stop_hook_active": false,
"last_assistant_message": "..."
}tool_use_id is the natural correlation key between a Pre and its matching Post invocation. The hook also has CLAUDE_PROJECT_DIR set as an env var (project-relative cwd for the running CC instance).
When PreToolUse returns hookSpecificOutput.updatedInput together with permissionDecision: "allow", the modified input takes effect for the actual tool execution but Claude is not informed of the change. Verified for Read with file_path redirect:
| Channel | Sees | Notes |
|---|---|---|
| Pre stdin | original input | Pre's chance to decide whether/how to modify |
| Tool execution | updated input | What CC actually runs |
Post stdin (tool_input) |
updated input | Reflects the post-modification state |
Post stdin (tool_response) |
result of running on updated input | e.g., content of the redirected file |
Transcript tool_use.input (Claude's record) |
original input | Claude believes it called what it asked for |
tool_result.content returned to Claude |
result of running on updated input | Claude sees the resulting bytes |
So a hook can transparently re-route or normalize a tool call without polluting Claude's mental model of what it did. The only signal Claude can pick up is content-level: if Claude reads the same path twice and sees different bytes, it can infer "something changed" but not "the path was redirected."
The empirical setup that produced this: PreToolUse(Read("orig.txt")) returning updatedInput.file_path = "redirect.txt" resulted in CC reading redirect.txt's content, Claude receiving that content, but Claude's transcript tool_use.input.file_path still showing orig.txt. Both Pre and Post hook stdin showed exactly what the columns above describe.
hookSpecificOutput.updatedInput (with permissionDecision: "allow") is supported since at least Claude Code 2.1.0. The 2.1.0 CHANGELOG entry is a fix for an ask-decision edge case, indicating the feature predates 2.1.0; the modern hookSpecificOutput schema has been stable since at least early 2.1.x. The plugin's existing reliance on ${CLAUDE_PLUGIN_ROOT} and stdin JSON shape already implies ≥ 2.1.0 so adding updatedInput use does not raise the floor in practice.
| Dimension | Edit |
Write |
|---|---|---|
| File existence | Must exist; CC also enforces "must have been Read" in the same session | May not exist (creates) or may exist (overwrites) |
| Modification mode | Surgical: replace old_string with new_string |
Full: replace entire content |
| Failure modes | old_string not found / matches multiple times |
Disk errors only (rare) |
| Typical input fields | {file_path, old_string, new_string} |
{file_path, content} |
For an encoding-handling plugin: Edit and Write on existing non-UTF-8 files share the same handling path (detect encoding, convert, do the operation, restore). Write on a non-existent file is a special case — there is no original encoding to detect, and the new file is naturally created as UTF-8.
Two architecturally distinct ways a hook plugin can keep the original bytes safe when a tool call fails:
(a) In-place convert + reverse (current main + chardet7-preview architecture): Pre converts the original file in place to UTF-8, Post converts back. If the tool fails without modifying content, the round-trip is byte-identical. Safety relies on a mathematical equivalence (encode(decode(bytes)) == bytes). Pre's convert_file aborts on UnicodeDecodeError so a wrong-encoding detection is caught before any write. The original file's mtime changes twice per Read/Edit cycle, which is visible to external watchers (IDEs, build daemons).
(b) Redirect to temp + sync-back on success (Solution 2 design, not yet implemented): Pre creates a fresh temp UTF-8 file and updatedInputs the path to it; the original is not touched. Post syncs back only on tool success; on failure, the temp is discarded and the original remains literally byte-identical with unchanged mtime. Safety is structural (Pre never writes the original) rather than mathematical. External watchers see at most one mtime bump per successful Edit/Write.
A failed tool — Edit with non-matching old_string, Write with disk error, etc. — leaves the original file's content intact under either architecture. Architecture (b) additionally preserves the mtime/inode-level invariants and gives stronger safety when Post itself encounters an error (e.g., reverse-encode failure due to characters Claude introduced that aren't representable in the original encoding — currently a known limitation in (a)).
These have not been verified by experiment yet (deferred until needed):
tool_responseschema forEdit/Writeon successtool_responseschema forEdit/Writeon failure (specifically, whether there is an explicit success/error field that hooks can branch on)- Pre/Post stdin
tool_inputshape forEdit(old_string/new_stringfield names) andWrite(content) - Whether
PostToolUsefires when the tool itself errors out (vs only on success) - Whether
Writeto a non-existent path triggers Pre with the path that doesn't exist yet (assumed yes)