#codecharts — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #codecharts, aggregated by home.social.
-
One would expect then that for these two blocks the names found in the NamesList.txt date file would be identical to the ones displayed in the code charts, but no, they are actually "combined":
@@ 0000 C0 Controls and Basic Latin (Basic Latin) 007F
@@ 0080 C1 Controls and Latin-1 Supplement (Latin-1 Supplement) 00FFBlocks.txt:
0000..007F; Basic Latin
0080..00FF; Latin-1 Supplement[ Pourquoi faire simple quand on peut faire compliqué ? ] (logique Shadok)
-
Section 24.1.13 of the Unicode 17.0 Core Spec states that:
"The page headers for the code charts are based on the normative values of the Block property defined in Blocks.txt in the Unicode Character Database, with a few exceptions." and indeed it seems that only the first two block names differ so far, possibly for legacy reasons (conformity with ISO?)."Basic Latin" !== "C0 Controls and Basic Latin"
"Latin-1 Supplement" !== "C1 Controls and Latin-1 Supplement" -
Two days ago, I reported issues about the French version of CodeCharts.pdf through the UTC Contact Form. Yesterday, I received a "reply" to my report:
- Totally unexpected and unnecessary reply.
- Apparently generated by AI.
- Verbose for no purpose.
- Off beam, riddled with typical hallucinations, mentioning things I never wrote.
- Not an "answer" to questions I never asked...#Unicode #GenerativeAI #Enshittification #CodeCharts #Report
-
Two days ago, I reported issues about the French version of CodeCharts.pdf through the UTC Contact Form. Yesterday, I received a "reply" to my report:
- Totally unexpected and unnecessary reply.
- Apparently generated by AI.
- Verbose for no purpose.
- Off beam, riddled with typical hallucinations, mentioning things I never wrote.
- Not an "answer" to questions I never asked...#Unicode #GenerativeAI #Enshittification #CodeCharts #Report
-
It seems that the Unicode code charts agree with you somehow ("preferred representation" is ambiguous though, it possibly just means which other character the glyph should be replicated from):
U+2126 OHM SIGN
* SI unit of resistance, named after G. S. Ohm, German physicist
* preferred representation is U+03A9
x (ascending node - U+260A)
: U+03A9 greek capital letter omega -
It seems that the Unicode code charts agree with you somehow ("preferred representation" is ambiguous though, it possibly just means which other character the glyph should be replicated from):
U+2126 OHM SIGN
* SI unit of resistance, named after G. S. Ohm, German physicist
* preferred representation is U+03A9
x (ascending node - U+260A)
: U+03A9 greek capital letter omega -
The #Unicode #codecharts are an amazing wealth of information. Apart from the representative glyphs which are nowhere else available, they can help to establish some kind of #taxonomy based on the names of #blocks and #subheaders, and also provide useful #annotations for #characters, creating #crossreferences between them.
🔗 https://www.unicode.org/Public/17.0.0/charts/CodeCharts.pdf
However, extracting this information from the NamesList.txt data file used to generate the charts proves to be uneasy…
-
The #Unicode #codecharts are an amazing wealth of information. Apart from the representative glyphs which are nowhere else available, they can help to establish some kind of #taxonomy based on the names of #blocks and #subheaders, and also provide useful #annotations for #characters, creating #crossreferences between them.
🔗 https://www.unicode.org/Public/17.0.0/charts/CodeCharts.pdf
However, extracting this information from the NamesList.txt data file used to generate the charts proves to be uneasy…
-
There is a entry in the #Unicode #CodeCharts which has been puzzling me since a while ago, both in the English and French version:
10F45 SOGDIAN INDEPENDENT SHIN
→ 6240 所10F45 CHINE INDÉPENDANT SOGDIEN
→ 6240 所Fortunately, I was able to find some explanation in the #CoreSpec - Chapter 14:
"The repertoire includes one phonogram, U+10F45 𐽅 SOGDIAN INDEPENDENT SHIN, an alternate form of isolated shin, used to transcribe one Chinese character, U+6240 所."
Now, I can die in peace…
-
There is a entry in the #Unicode #CodeCharts which has been puzzling me since a while ago, both in the English and French version:
10F45 SOGDIAN INDEPENDENT SHIN
→ 6240 所10F45 CHINE INDÉPENDANT SOGDIEN
→ 6240 所Fortunately, I was able to find some explanation in the #CoreSpec - Chapter 14:
"The repertoire includes one phonogram, U+10F45 𐽅 SOGDIAN INDEPENDENT SHIN, an alternate form of isolated shin, used to transcribe one Chinese character, U+6240 所."
Now, I can die in peace…
-
I just found incidentally this interesting document about Unicode code charts; still a draft, I believe it is new in Unicode 18.0...
🔗 https://www.unicode.org/Public/draft/charts/About.html
And of course, there is chapter 24 of the outstanding "Core Spec": About the Code Charts.
🔗 https://www.unicode.org/versions/Unicode18.0.0/core-spec/chapter-24/
-
I just found incidentally this interesting document about Unicode code charts; still a draft, I believe it is new in Unicode 18.0...
🔗 https://www.unicode.org/Public/draft/charts/About.html
And of course, there is chapter 24 of the outstanding "Core Spec": About the Code Charts.
🔗 https://www.unicode.org/versions/Unicode18.0.0/core-spec/chapter-24/
-
Il existe une version en français des "Code Charts Unicode", dont peu de gens soupçonnent même l'existence...
Français: https://www.unicode.org/Public/17.0.0/charts/fr/CodeCharts.pdf
(Anglais: https://www.unicode.org/Public/17.0.0/charts/CodeCharts.pdf)Aujourd'hui, je viens de trouver, un peu par hasard, la version française ListeNoms.txt (apparemment québécoise) du fichier NamesList.txt utilisé justement pour générer les données des "code charts":
Français: https://hapax.qc.ca/ListeNoms-17.0.0.txt
(Anglais: https://www.unicode.org/Public/17.0.0/ucd/NamesList.txt) -
Il existe une version en français des "Code Charts Unicode", dont peu de gens soupçonnent même l'existence...
Français: https://www.unicode.org/Public/17.0.0/charts/fr/CodeCharts.pdf
(Anglais: https://www.unicode.org/Public/17.0.0/charts/CodeCharts.pdf)Aujourd'hui, je viens de trouver, un peu par hasard, la version française ListeNoms.txt (apparemment québécoise) du fichier NamesList.txt utilisé justement pour générer les données des "code charts":
Français: https://hapax.qc.ca/ListeNoms-17.0.0.txt
(Anglais: https://www.unicode.org/Public/17.0.0/ucd/NamesList.txt) -
Hype for the Future 49U: Additional Interests with Unicode
Thanks to Unicode, nearly every language and writing system across the globe can be represented online. Even emojis have designated sections of the code charts, along with the tags of the associated emoji codes. Unfortunately, however, subnational flags may be re-rendered as black flags at times, since the tags in Unicode may not be maintained as appropriately as such. England, Scotland, and Wales are examples of such subnational entity flags. Beyond the writing systems and the obvious, even […] -
Hype for the Future 49T: What about Romanian?
The Romanian language requires the comma accents below the letters S and T as accented letters that change the pronunciation of certain words in the Romanian language. Even though specifically the comma letters are in Latin Extended-B, many of the other Romanian accents exist in Latin Extended-A or perhaps even in Latin-1 Supplement. Therefore, the Latin Extended-B characters are specifically designated as Romanian additions. Otherwise, most of the letters in Latin Extended-B and later down […]https://novatopflex.wordpress.com/2025/12/19/hype-for-the-future-49t-what-about-romanian/
-
Hype for the Future 49T: What about Romanian?
The Romanian language requires the comma accents below the letters S and T as accented letters that change the pronunciation of certain words in the Romanian language. Even though specifically the comma letters are in Latin Extended-B, many of the other Romanian accents exist in Latin Extended-A or perhaps even in Latin-1 Supplement. Therefore, the Latin Extended-B characters are specifically designated as Romanian additions. Otherwise, most of the letters in Latin Extended-B and later down […]https://novatopflex.wordpress.com/2025/12/19/hype-for-the-future-49t-what-about-romanian/
-
Hype for the Future 49R: Latin Extended-A
Spanning U+0100 and U+017F of the Unicode Code Charts, Latin Extended-A predominantly services localization and internationalization for Eastern European languages in the Latin script, including Polish, Latvian, Lithuanian, Turkish, Czech, and Slovak, though not every accented letter may necessarily be represented in the code block for any of the aforementioned languages.
-
From time to time (since this represents a tremendous amount of translation/adaptation work), a French version of the "code charts" gets published by the Unicode Consortium: the latest one is for Unicode 16.0:
https://www.unicode.org/Public/16.0.0/charts/fr/CodeCharts.pdf
This is especially useful for French speakers in #Canada, #France, #Belgium, #Switzerland, etc. but may soon be obsolete for #Quebec, in case it gets "absorbed" by a neighboring country whose official language is now English only...
-
De temps en temps (cela représente un énorme travail d'adaptation), une version française des "code charts" est publiée par le Consortium Unicode, la dernière en date est pour Unicode 16.0:
https://www.unicode.org/Public/16.0.0/charts/fr/CodeCharts.pdf
Malheureusement, celle-ci risque d'être bientôt obsolète pour les francophones de la belle province de Québec, dans le cas où celle-ci serait «absorbée» par un pays voisin dont la langue officielle est désormais uniquement l'anglais...
-
Unicopedia Anatolica is a developer-oriented set of #Unicode utilities related to Anatolian hieroglyphs, wrapped into one single app, built with #Electron.
Repository: 🔗 https://codeberg.org/tonton-pixel/unicopedia-anatolica
#anatolian #hieroglyphs #unicopedia #javascript #unicode #characters #codepoints #codecharts #desktopapplication #electronjs #glyphs #localfonts
-
Unicopedia Ægypta is a developer-oriented set of #Unicode utilities related to Egyptian hieroglyphs, wrapped into one single app, built with #Electron.
Repository: 🔗 https://codeberg.org/tonton-pixel/unicopedia-aegypta
#characters #codecharts #codepoints #desktopapplication #egyptian #electronjs #glyphs #hieroglyph #hieroglyphs #javascript #localfonts #unicode #unicopedia #unikemet
-
Unicopedia Sinica is a developer-oriented set of #Unicode utilities related to ideographs, wrapped into one single app, built with #Electron.
Repository: 🔗 https://codeberg.org/tonton-pixel/unicopedia-sinica
#characters #chinese #cjk #cjkrelated #cjkv #codecharts #codepoints #components #confusables #desktopapplication #electronjs #glyphs #ideographs #ideographicdescriptionsequences #ids #japanese #javascript #kangxi #kangxiradicals #korean #localfonts #opensource #strokes #tangut #unicode #unicopedia #unihan #vietnamese