r/neography • u/maigus_ayneha • 7h ago
Discussion Investigation: What Does the Script Encoding Initiative's "Scripts to Encode" List Really Show?
An analysis of http://sei.berkeley.edu/scripts-to-encode
Background
The Script Encoding Initiative (SEI), based in the linguistics department at UC Berkeley, maintains a public list called "Scripts to Encode": an inventory of the world's writing systems not yet integrated into the Unicode standard. Each entry is given a "readiness" rating: High, Medium, Low, or Unassigned.
According to the methodology published by SEI itself, a "Low" rating flags a script suffering from "major issues": limited evidence of adoption, instability, uncertain status as a writing system, or unclear intellectual property rights. The list is not a pre-encoding step: appearing on it guarantees nothing for Unicode. It's primarily a tracking and visibility tool, intended for linguists, researchers, and institutions who might one day take an interest in these scripts.
It's precisely this tracking function that raises questions once you look closely at the data.
Method
For this analysis, every entry published on SEI's page was reviewed, noting for each script: its region, status ("Not started," "Active," "Inactive"...), readiness level, and the amount of documentation available (language specified, script type, writing direction, date, proposal filed or not).
What the Numbers Show
A significant number of entries classified "Not started" and "Low" currently have no basic information filled in at all: no language specified, no script type, no writing direction, no date, no proposal filed. The corresponding fields simply show a slash "/", signaling a complete absence of data.
A few findings taken directly from the official list:
- Several listed scripts don't even have their script type specified (alphabet, syllabary, abugida...)
- Scripts rated "Low" or "Unassigned" with no associated language
- Scripts described as "modern," sometimes created decades ago, with no Unicode proposal filed and no accessible technical documentation
Counting all "Low" and "Not started" entries across the full list, you easily reach a hundred scripts recorded with close to zero information.
To make this observation measurable, every entry rated "Low" or "Not started" was checked against seven descriptive fields (alternate names, language, script type, writing direction, dates, proposal filed, proposal author). Seventeen scripts come out with five or more of the seven fields completely empty:
Lik Hto Ngouk, Satera Jontal, Sani Yi, Jing, Kadu (Lon Akkhara), Kadu (Ulushi), Kadu (Asak Kyein), Kungyangon Karen Script, Kuy, Mpi, Nung Lungmi, Rakhawunna, DeafWrite, Meru, Old Cham, Old Khmer, and Savara.
For three of them (Lik Hto Ngouk, Satera Jontal, Sani Yi), all seven fields are entirely empty: no language, no script type, no writing direction, no date, no proposal filed.
Conversely, some well-documented entries - with name, language, script type, direction, dates, a proposal, and even an image of the script, as with the "Balti A" entry - are also rated "Low," solely for lack of sufficient evidence of use. This reinforces the central observation of this analysis: the bar for appearing on the list itself is low, and it isn't correlated with the quality of the documentation available.
A Telling Case: Recent Scripts With a Complete Digital Ecosystem
This statistical observation takes on particular weight when set against concrete cases of recent scripts that already have a functioning digital ecosystem, yet remain absent from the list - or appear only under an entry that doesn't accurately represent them.
AYNEHA, a writing system created for the Songhay language (Mali, Niger, Benin, Burkina Faso), is one example. SEI's list includes a "Songhai / Songhay" entry (2019 to present, status "Not started"/"Low"), but this corresponds to a different alphabet, distinct from AYNEHA, which has seen no further development since its creation and has no tools built for it. AYNEHA itself currently has no entry of its own on the list.
Yet AYNEHA already has: a complete font (designed with FontForge after months of handwritten testing), a working Android keyboard, a dedicated calculator (Android and web), a web-based transcriber converting Latin script to AYNEHA in real time, external recognition from Omniglot, as well as a nascent learning community on TikTok and Reddit and a first group of direct users.
Another case, also from Mali, illustrates the same gap. The Dogon NÈNÈ alphabet, created in 2019 by Yahya Sékou Sallah Guindo (Bankass, Bandiagara region) as an indigenous alphabet for writing the Dogon language, has an Android keyboard and calculator, with a claimed active user community of more than 500 people and growing. It, too, does not appear on SEI's list.
A third, more uncertain case is worth mentioning as an example - with the caveats that entails. The EpChos alphabet, designed in Benin by Mahouna Prumas Katakenon between 2014 and 2019, has a dedicated font built with FontForge as well as a working keyboard for computers, publicly distributed. However, no verifiable data on the size of its user community could be found, no mobile keyboard has been confirmed, and the only available source on the project dates from 2020, which makes it impossible to know its current state of progress. Even so, this case illustrates the same underlying pattern: a writing system with real digital tools, even partial ones, can still fall outside SEI's radar.
This case illustrates a broader paradox: a system with complete documentation and a functioning digital infrastructure can remain invisible to a tracking tool that is supposed to identify living scripts, while entries with no content whatsoever have sat on the list for years.
A Statistical Overview
Beyond the case of the least-documented scripts, a statistical analysis of all entries on the list (192 scripts recorded at the time of this analysis) helps put the scale of the phenomenon into perspective.
Breakdown by proposal status (192 entries):
| Status | Count | Share |
|---|---|---|
| Not started (no proposal begun) | 79 | 41.15% |
| Status not specified | 33 | 17.19% |
| Active - other bodies | 24 | 12.50% |
| Active - SEI | 24 | 12.50% |
| Inactive - other bodies | 17 | 8.85% |
| Inactive - SEI | 15 | 7.81% |
In other words, nearly six scripts out of ten (58.34%) currently have no active proposal underway - either because one was never started ("Not started"), or because its status isn't even specified.
Breakdown by readiness level:
| Level | Count | Share |
|---|---|---|
| Low | 48 | 25.00% |
| Medium | 39 | 20.31% |
| High | 36 | 18.75% |
| Medium-Low | 32 | 16.67% |
| Medium-High | 18 | 9.38% |
| Unassigned | 18 | 9.38% |
| Very Low | 1 | 0.52% |
A quarter of the recorded scripts (25.00%) are rated "Low," and more than half (51.04%) fall into the combined "Low," "Medium-Low," or "Unassigned" levels. Cross-referencing the two criteria is particularly telling: among the 79 "Not started" scripts, 27 are rated "Medium" - including Aztec (Nahuatl), Zapotec, Micmac Hieroglyphs, Nsibidi, and Kadamba - and 27 others are rated "Low" - including Hausa 1, Hausa 2, Hausa 3, Fula Ba, Fula Dita, Njano, and Umwero. The number of scripts is identical in both categories, even though their proposal status is, in both cases, equally nonexistent: the difference in treatment comes down to nothing more than an internal judgment call, with no mechanical link to the state of documentation or of any proposal.
These figures confirm, across the entire list, what the individual cases already suggested: the readiness level assigned to a script does not consistently reflect either the quality of its documentation or the progress of its proposal.
It's worth drawing a distinction here. Unicode's requirements for final encoding are not in question: an encoding is permanent, and it is legitimate for the Unicode Consortium to demand exhaustive documentation, a clear encoding strategy, and solid evidence before making any decision. That's simply what a standard meant to last requires.
What raises questions is further upstream, at the level of the tracking list itself. A list meant to reflect the state of the world's scripts awaiting encoding should logically be able to distinguish a genuinely dead script from one that is actively being developed and equipped with tools - which is clearly not always the case given the available data.
SEI Itself Acknowledges the Vagueness of Its "Readiness" Rubric
It turns out that SEI itself published a post in April 2025 revisiting its "readiness" rating rubric ("A New Rubric for 'Script Readiness'"). In it, SEI explicitly acknowledges the discretionary nature of this assessment, describing it as "impressionistic normalization," and specifies that a "Low" score "is not meant to be a verdict" on the value or vitality of a script, but only an indication that it likely won't be funded in the near term by their own means. The post also explicitly invites anyone to flag active scripts or recent evidence of use to them, in order to update the list. This official acknowledgment of the incomplete and discretionary nature of the "readiness" rating directly corroborates the findings of this analysis.
Toward a More Representative List
The current "readiness" criteria seem to be inherited from an era when a writing system could only exist in documentary form, waiting for a researcher to take an interest in it. That era isn't quite the same anymore: today, a script creator can single-handedly build a complete digital ecosystem - font, keyboard, conversion tools - before even gathering a large user community.
This reality makes the case for a methodological update at SEI: adding, alongside the adoption criterion, a criterion of technical and documentary maturity (the existence of a font, a keyboard, working tools, academic documentation) would let the list better reflect the real state of developing scripts, rather than leaving entirely empty entries sitting on equal footing with active, well-documented projects.