Important
This repository is intentionally a TypeScript port and derived work of
Robyn Speer's rspeer/python-ftfy.
Its algorithms, behavior, tests, examples, and documentation are based on
upstream commit
74dd0452b48286a3770013b3a02755313bd5575e.
This project does not claim independent authorship of ftfy. See
Attribution and license and UPSTREAM.md.
import { fix_encoding } from "ftfy-js";
console.log(fix_encoding("(ง'⌣')ง"));
// (ง'⌣')งftfy fixes Unicode text that is broken in various ways. Its main goal is to
take in bad Unicode and output good Unicode, particularly by repairing
mojibake: text that was decoded using the wrong character encoding.
This package mirrors python-ftfy 6.3.1's Python-style API names so behavior and
examples map directly between the two projects. It ships ESM, CommonJS,
TypeScript declarations, browser-compatible core functions, and a Node.js CLI.
It is published to npm as ftfy-js.
npm install ftfy-jsFor local development:
npm install
npm testNode.js 20 or newer is supported.
ftfy detects character patterns that were probably intended to be UTF-8 but were decoded as another encoding:
import { fix_text } from "ftfy-js";
fix_text("✔ No problems");
// "✔ No problems"It can repair multiple layers of mojibake:
fix_text("The Mona Lisa doesn’t have eyebrows.");
// "The Mona Lisa doesn't have eyebrows."It can fix mojibake after curly quotes have been applied:
fix_text("l’humanité");
// "l'humanité"It can recover a non-breaking-space byte that was changed into an ASCII space:
fix_text("à perturber la réflexion");
// "à perturber la réflexion"
fix_text("à perturber la réflexion");
// "à perturber la réflexion"It can decode HTML entities outside HTML, including carefully selected incorrectly-capitalized forms:
fix_text("PÉREZ");
// "PÉREZ"The fixes are deliberately conservative. Correctly decoded text that merely could be interpreted as mojibake is left alone:
fix_text("IL Y MARQUÉ…");
// "IL Y MARQUÉ…"The common functions are:
import {
TextFixerConfig,
apply_plan,
fix_and_explain,
fix_encoding,
fix_encoding_and_explain,
fix_text,
fix_text_segment,
guess_bytes,
} from "ftfy-js";fix_text(text, config?)applies all enabled fixes, independently by line.fix_text_segment(text, config?)applies a single consistent sequence of fixes to the whole input.fix_encoding(text, config?)repairs encoding mix-ups only.fix_and_explain(text, config?)andfix_encoding_and_explain(text, config?)return an array that can be destructured as[text, explanation]and also has.textand.explanationproperties.apply_plan(text, explanation)replays an explanation.guess_bytes(bytes)implements upstream's deliberately limited byte guesser.TextFixerConfigexposes the same snake_case settings as python-ftfy.
const [fixed, explanation] = fix_and_explain("só");
fixed;
// "só"
explanation;
// [["encode", "latin-1"], ["decode", "utf-8"]]
apply_plan("só", explanation ?? []);
// "só"Configuration can be passed as a plain object or as a TextFixerConfig:
fix_text("“text”", { uncurl_quotes: false });
const config = new TextFixerConfig({
unescape_html: false,
normalization: "NFKC",
});
fix_text("text", config);Individual fixers are exported from ftfy-js/fixes; heuristic functions from
ftfy-js/badness; encoding data from ftfy-js/chardata; terminal-width helpers
from ftfy-js/formatting; and codec helpers from ftfy-js/bad-codecs.
See the full documentation.
The package provides the same ftfy command-line shape as upstream:
ftfy input.txt
ftfy -e latin-1 input.txt -o output.txt
ftfy -g input.txt
cat input.txt | ftfyRun ftfy --help for all options. The CLI reads and writes UTF-8 by default.
The port is pinned to the upstream revision documented in UPSTREAM.md. Its test suite directly ports all upstream unit tests and all five upstream JSON corpora. Known upstream failures remain explicit tests and are not presented as successes. The codec, HTML entity, Unicode-name, Unicode-category, and display width tables are generated from the pinned upstream Python environment instead of being approximated with unrelated JavaScript packages.
Run the complete release check with:
npm run checkThis runs strict TypeScript checking, all tests, both ESM and CommonJS builds, the CLI build, and an npm package dry run. npm publishing is intentionally not part of any script or workflow.
ftfy was designed and created by Robyn Speer, also known as Elia Robyn
Lake. The canonical project is
rspeer/python-ftfy. Robyn's projects
and writing are available on GitHub and at
posts.arborelia.net.
This TypeScript repository is a language port of that work.
ftfy has been used as a crucial data-processing step in major NLP research. When using this port in research, cite the original work:
Robyn Speer. (2019). ftfy (Version 5.5). Zenodo.
https://doi.org/10.5281/zenodo.2591652
@misc{speer-2019-ftfy,
author = {Robyn Speer},
title = {ftfy},
note = {Version 5.5},
year = 2019,
howpublished = {Zenodo},
doi = {10.5281/zenodo.2591652},
url = {https://doi.org/10.5281/zenodo.2591652}
}This repository follows the Important license clarifications in python-ftfy's README.
The Apache License 2.0 is what grants permission to use, copy, modify, and distribute ftfy. Following that license is required. In particular, derived works must retain the required notices and must correctly attribute ftfy's author, Robyn Speer. A derived work must not obscure ftfy's authorship.
This repository therefore:
- identifies itself prominently as a TypeScript port and derived work;
- names Robyn Speer as the original author;
- links to the canonical python-ftfy repository;
- records the exact upstream commit;
- reproduces the upstream Apache-2.0 license in LICENSE.txt; and
- includes attribution in NOTICE, generated files, tests, and docs.
See the upstream Important license clarifications for the author's complete statement. If this project is redistributed or further modified, do not remove or obscure this attribution.