home.social

#unicode — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #unicode, aggregated by home.social.

fetched live
  1. Почему кросс-постинг в соцсети оказался задачей про state, idempotency и битую кириллицу

    В какой-то момент у меня появилась задача от маркетинга автоматизировать кросс-постинг статей в соц.сети и, в принципе, задача довольно понятная и здравая. Запрос был такой - сделать короткий анонс из готовой статьи в блоге для размещения в соц.сети, подготовить рерайты для площадок, прикрепить картинку, поставить UTM-метки и опубликовать все по каналам. Т.к. задача рутинная, сразу возникла идея отдать ее агенту. То есть, на входе статья в блоге, а на выходе Google Doc с короткими анонсами, рерайтами для Дзена и Spark, а далее публикация в Telegram и отложенные посты для остальных социальных сетей. При первом же нормальном прогоне стало ясно что в этой задаче много подводных камней и заработало это все только когда накопившиеся ошибки начали превращаться в инварианты, проверки и стоп-факторы. Короче, распишу все эти проблемы, может кому пригодится в работе, учитывая что маркетинг сейчас очень сильно хочет в ИИ.

    habr.com/ru/articles/1070282/

    #ai #автоматизация_процессов #социальные_сети #workflow #api #devops #интеграция_сервисов #качество_данных #unicode #google_docs

  2. Here's one for the encoding-bi-curious: Is there a UTF-8 <=> ASCII/Windows-1252 interchange encoding?

    Imagine you're writing a text-editor for an operating system that predates UTF-8 (I don't have to imagine), and you can't do anything about what the OS stores and renders as text. It's hard-coded to Windows-1252[^1] and you're regularly transferring files to/from PC that are utf-8 encoded. What do you do?

    The format in RAM *must* be CP1252, but it can be serialised to/from disk as UTF-8. The first part is easy, just convert any special characters that *are* available in CP1252 (e.g. "÷") to preserve those. But what about characters that have no representation? (see images)

    I could convert graphemes to X/HTML entities (`&copy;`) on load but then when saving it would convert pre-existing X/HTML entities into UTF-8. Same goes with other encodings like C/C++ escape-sequences (`\x1B`) but I suppose that depends on what type of text you're editing.

    [^1]: en.wikipedia.org/wiki/Windows-

    #psion #programming #unicode

  3. Here's one for the encoding-bi-curious: Is there a UTF-8 <=> ASCII/Windows-1252 interchange encoding?

    Imagine you're writing a text-editor for an operating system that predates UTF-8 (I don't have to imagine), and you can't do anything about what the OS stores and renders as text. It's hard-coded to Windows-1252[^1] and you're regularly transferring files to/from PC that are utf-8 encoded. What do you do?

    The format in RAM *must* be CP1252, but it can be serialised to/from disk as UTF-8. The first part is easy, just convert any special characters that *are* available in CP1252 (e.g. "÷") to preserve those. But what about characters that have no representation? (see images)

    I could convert graphemes to X/HTML entities (`&copy;`) on load but then when saving it would convert pre-existing X/HTML entities into UTF-8. Same goes with other encodings like C/C++ escape-sequences (`\x1B`) but I suppose that depends on what type of text you're editing.

    [^1]: en.wikipedia.org/wiki/Windows-

    #psion #programming #unicode

  4. Retour d'un troll vieux de plus de 10 ans
    Le Consortium Unicode va-t-il résister ?

    Pour DJ Snake, « le drapeau breton mérite son émoji » : la star relance un vieux débat sur les réseaux sociaux
    ouest-france.fr/bretagne/pour-
    #bretagne #unicode #utf8

  5. Retour d'un troll vieux de plus de 10 ans
    Le Consortium Unicode va-t-il résister ?

    Pour DJ Snake, « le drapeau breton mérite son émoji » : la star relance un vieux débat sur les réseaux sociaux
    ouest-france.fr/bretagne/pour-
    #bretagne #unicode #utf8

  6. 🐔✨ Chicken Scheme 6.0: Now with 100% more #Unicode and #bytevectors, because obviously the world can't function without #UTF8 chicken blobs being dethroned by bytevectors. Who needs coherent syntax when you can have charname and u8... discussions during a dinner party? 🙃💾
    code.call-cc.org/releases/6.0. #ChickenScheme #CodingDiscussions #HackerNews #ngated

  7. 🐔✨ Chicken Scheme 6.0: Now with 100% more #Unicode and #bytevectors, because obviously the world can't function without #UTF8 chicken blobs being dethroned by bytevectors. Who needs coherent syntax when you can have charname and u8... discussions during a dinner party? 🙃💾
    code.call-cc.org/releases/6.0. #ChickenScheme #CodingDiscussions #HackerNews #ngated

  8. @gloriouscow @foone
    If you strategize carefully and maybe had insider help you might do an #emoji proposal or two targeted at making #Doom happen

    I believe all the code points are there waiting to be used

    They just don’t present as emoji (or any other character) yet

    So if you know what you want you could build towards that

    You can post unreleased emoji and they show up as boxes until your OS is updated to display them

    DOS might not care, but unsatisfying until #Unicode approves them

  9. @gloriouscow @foone
    If you strategize carefully and maybe had insider help you might do an #emoji proposal or two targeted at making #Doom happen

    I believe all the code points are there waiting to be used

    They just don’t present as emoji (or any other character) yet

    So if you know what you want you could build towards that

    You can post unreleased emoji and they show up as boxes until your OS is updated to display them

    DOS might not care, but unsatisfying until #Unicode approves them

  10. 🚀 Ah, TSON: because #JSON just wasn't complex enough! Now with hash-pinned metaschemas for the masochists who think #Unicode is a data format 🍜. Enjoy writing "data" that's solely understandable to machines and #AI enthusiasts living in 2099 🌌.
    tson.io/ #TSON #Complexity #DataFormat #HackerNews #ngated

  11. 🚀 Ah, TSON: because #JSON just wasn't complex enough! Now with hash-pinned metaschemas for the masochists who think #Unicode is a data format 🍜. Enjoy writing "data" that's solely understandable to machines and #AI enthusiasts living in 2099 🌌.
    tson.io/ #TSON #Complexity #DataFormat #HackerNews #ngated

  12. One would expect then that for these two blocks the names found in the NamesList.txt date file would be identical to the ones displayed in the code charts, but no, they are actually "combined":

    @@ 0000 C0 Controls and Basic Latin (Basic Latin) 007F
    @@ 0080 C1 Controls and Latin-1 Supplement (Latin-1 Supplement) 00FF

    Blocks.txt:

    0000..007F; Basic Latin
    0080..00FF; Latin-1 Supplement

    [ Pourquoi faire simple quand on peut faire compliqué ? ] (logique Shadok)

    #Unicode #Blocks #CodeCharts

  13. Section 24.1.13 of the Unicode 17.0 Core Spec states that:
    "The page headers for the code charts are based on the normative values of the Block property defined in Blocks.txt in the Unicode Character Database, with a few exceptions." and indeed it seems that only the first two block names differ so far, possibly for legacy reasons (conformity with ISO?).

    "Basic Latin" !== "C0 Controls and Basic Latin"
    "Latin-1 Supplement" !== "C1 Controls and Latin-1 Supplement"

    #Unicode #CodeCharts #CoreSpec

  14. Minor rant about #Unicode. Simply put, text handling sucks. I get that we no longer live in an ASCII world, that we have to handle much more than that. But the tools we have don't handle that particularly well at this time.

    It's text. It shouldn't be so easy to break it apart and render things unreadable. And yet that is exactly what I have to contend with. Silly little bugs that break apart a single Unicode character just because it happens to be more than one byte long.

    There has to be something better. There just has to be.

  15. Minor rant about #Unicode. Simply put, text handling sucks. I get that we no longer live in an ASCII world, that we have to handle much more than that. But the tools we have don't handle that particularly well at this time.

    It's text. It shouldn't be so easy to break it apart and render things unreadable. And yet that is exactly what I have to contend with. Silly little bugs that break apart a single Unicode character just because it happens to be more than one byte long.

    There has to be something better. There just has to be.

  16. @wyatt Changing character set for blocks of text is what Standard Compression Scheme for Unicode (SCSU) tried to do. However, because it allows changing to UTF-16BE for blocks of CJK text, this compressing character encoding is not a superset of ASCII and therefore not suitable for things like HTML because of a cross-site scripting weakness.
    en.wikipedia.org/wiki/Standard

    #Unicode #SCSU

  17. @wyatt Changing character set for blocks of text is what Standard Compression Scheme for Unicode (SCSU) tried to do. However, because it allows changing to UTF-16BE for blocks of CJK text, this compressing character encoding is not a superset of ASCII and therefore not suitable for things like HTML because of a cross-site scripting weakness.
    en.wikipedia.org/wiki/Standard

    #Unicode #SCSU

  18. unicode consortium HAS TO ADD head-ruffle or head-pat emoji

    #emoji #unicode

  19. I know someone out there in the blind community is going to ask why, so here's the reason: If you choose to use the setting in the language settings in Windows 11 25H2 labeled "Use Unicode (UTF-8) for worldwide language support" and have installed all the supplemental fonts that Windows supports in the Optional Features settings page (a long process to install), using certain TTS engines in NVDA will result in weird text in places. The only TTS that seems not to have an issue is eSpeak-NG. No weird text, just proper language support. What this means for me is that my favorite song titles and band/artist names in the German language render properly instead of looking odd. #A11Y #Accessibility #Blind #NVDA #eSpeakNG #Unicode #UTF8

  20. I know someone out there in the blind community is going to ask why, so here's the reason: If you choose to use the setting in the language settings in Windows 11 25H2 labeled "Use Unicode (UTF-8) for worldwide language support" and have installed all the supplemental fonts that Windows supports in the Optional Features settings page (a long process to install), using certain TTS engines in NVDA will result in weird text in places. The only TTS that seems not to have an issue is eSpeak-NG. No weird text, just proper language support. What this means for me is that my favorite song titles and band/artist names in the German language render properly instead of looking odd. #A11Y #Accessibility #Blind #NVDA #eSpeakNG #Unicode #UTF8

  21. #TIL that in some 12 based number systems (i.e. #duodecimal systems), the digits for 10 and 11 are not A and B as common in hexadecimal systems, but a turned 2 and 3 digit:

    ↊ — U+218A TURNED DIGIT TWO

    ↋ — U+218B TURNED DIGIT THREE

    #Unicode

  22. #TIL that in some 12 based number systems (i.e. #duodecimal systems), the digits for 10 and 11 are not A and B as common in hexadecimal systems, but a turned 2 and 3 digit:

    ↊ — U+218A TURNED DIGIT TWO

    ↋ — U+218B TURNED DIGIT THREE

    #Unicode

  23. Le visage craquelé remplace le visage aux yeux plissés dans les emojis de 2027 dlvr.it/TTctzk #Emojis #Unicode

  24. Le visage craquelé remplace le visage aux yeux plissés dans les emojis de 2027 dlvr.it/TTctzk #Emojis #Unicode

  25. Late-night icon picker for terminal people 😍

    🌃 **latuicon** — A TUI icon picker for emoji, kaomoji, Unicode & Nerd Font glyphs

    💯 Fuzzy search, multiple icon sets, themes, mouse support & stdout-friendly selection

    🦀 Written in Rust & built with @ratatui_rs

    ⭐ GitHub: github.com/coko7/latuicon

  26. Late-night icon picker for terminal people 😍

    🌃 **latuicon** — A TUI icon picker for emoji, kaomoji, Unicode & Nerd Font glyphs

    💯 Fuzzy search, multiple icon sets, themes, mouse support & stdout-friendly selection

    🦀 Written in Rust & built with @ratatui_rs

    ⭐ GitHub: github.com/coko7/latuicon

    #rustlang #tui #ratatui #unicode #nerdfont #emoji #terminal

  27. World Emoji Day is almost over. Why did I build Emoji Anki?

    As often, it was a silly idea. I thought that learning language vocabulary with emojis would be fun. I was already using Ankidroid for spaced repetition; there were a few Anki decks with emojis, but most were incomplete, or did not allow chosing categories, there was (very) limited language choice.

    Emoji Anki uses the publicly available Unicode CLDR data, that has annotations for every emojis, in 150+ languages.

    As a side effect, another fun thing you can do with it is learn country flags.

    anki.emoji-on.top/

    #EmojiAnki #WorldEmojiDay #Unicode #Emojis #LanguageLearning #Anki

  28. World Emoji Day is almost over. Why did I build Emoji Anki?

    As often, it was a silly idea. I thought that learning language vocabulary with emojis would be fun. I was already using Ankidroid for spaced repetition; there were a few Anki decks with emojis, but most were incomplete, or did not allow chosing categories, there was (very) limited language choice.

    Emoji Anki uses the publicly available Unicode CLDR data, that has annotations for every emojis, in 150+ languages.

    As a side effect, another fun thing you can do with it is learn country flags.

    anki.emoji-on.top/

    #EmojiAnki #WorldEmojiDay #Unicode #Emojis #LanguageLearning #Anki

  29. 17 July is the one day of the year when the calendar #emoji is correct!

    The word emoji combines the Japanese e (絵, picture) + moji (文字, character) and dates back to the late 80s. Its similarity to the English "emotion" is just a coincidence.

    Nowadays, emoji are standardised by the Unicode Consortium, a Californian non-profit that has its roots in facilitating a universal digital representation of the world's writing systems. #Unicode doesn't specify exactly how emoji should look – so they're different on Android and Apple devices, for example – but it lists their names.

    The first major character set used in computing was #ASCII, the American Standard Code for Information Interchange, published in 1963. Using 7 bits, it could represent 128 characters — more than enough for a to z, A to Z, 0 to 9 and common punctuation marks, but certainly indicative of a North American worldview.

    Later "extended" versions of ASCII doubled this to 256 possibilities, allowing for dozens of accented letters, characters like æ and ß, and symbols like © and ¾. And double-byte character sets like Shift JIS meant that Chinese, Korean and Japanese could be represented in full. But multiple standards existed, and it became just a bit too common to see ����.

    The answer was Unicode: a single, multi-byte character encoding standard for everyone. Plus emoji!

    #language #i18n

  30. Happy World Emoji Day! I was so preoccupied with whether I could, I did not think whether I should.

    https://ḧ̴͖́e̷͚̿-̸̧͘c̴͖͌o̴̻̊m̷͕̂e̷͔͊t̷͚̊h̵̦̄.op-co.de/

    Uncursed URL: op-co.de/blog/posts/emoji_xmpp

    Three weeks in the making. 22 years of standardization. 19 RFCs. 17 Unicode standards. 5000 words.

    cc @zeank @little4744 @chewie @M0YNG @hook

    #SorryNotSorry #xmpp #Emoji #IDNA #PRECIS #Unicode #Zalgo

  31. Happy World Emoji Day! I was so preoccupied with whether I could, I did not think whether I should.

    https://ḧ̴͖́e̷͚̿-̸̧͘c̴͖͌o̴̻̊m̷͕̂e̷͔͊t̷͚̊h̵̦̄.op-co.de/

    Uncursed URL: op-co.de/blog/posts/emoji_xmpp

    Three weeks in the making. 22 years of standardization. 19 RFCs. 17 Unicode standards. 5000 words.

    cc @zeank @little4744 @chewie @M0YNG @hook

    #SorryNotSorry #xmpp #Emoji #IDNA #PRECIS #Unicode #Zalgo

  32. Introducing Aksharapad — a free, browser-based editor for Sanskrit and other Indic languages.

    It combines a Unicode text editor with built-in transliteration, making it easy to write in Devanagari directly in your browser. No installation, no account, and it works offline after the first load.

    Try it here:
    oleandrum.github.io/aksharapad/

    Feedback, suggestions, and contributions are always welcome.

    #Sanskrit #Devanagari #Unicode #Indic #OpenSource #JavaScript #Writing #Linguistics

  33. Introducing Aksharapad — a free, browser-based editor for Sanskrit and other Indic languages.

    It combines a Unicode text editor with built-in transliteration, making it easy to write in Devanagari directly in your browser. No installation, no account, and it works offline after the first load.

    Try it here:
    oleandrum.github.io/aksharapad/

    Feedback, suggestions, and contributions are always welcome.

    #Sanskrit #Devanagari #Unicode #Indic #OpenSource #JavaScript #Writing #Linguistics

  34. I created Lipyantara — a free, browser-based tool for transliterating between Devanagari and Latin transliteration systems.

    It supports Sanskrit with a focus on accuracy, speed, and offline use. No installation, no account required.

    Try it here:
    oleandrum.github.io/lipyantara/

    Feedback, bug reports, and contributions are always welcome.

    #Sanskrit #Devanagari #Unicode #Transliteration #Indic #OpenSource #JavaScript #Linguistics

  35. Unicode’s UTS #35 transliteration rules are Turing-complete, revealing hidden computing power and new security risks inside the ICU library. hackernoon.com/unicode-transli #unicode

  36. RE: scicomm.xyz/@superlinguo/11692

    Another year, another #emoji release without a #cymbal. We’ll have two butterflies to choose from, a meteor AND a comet to ensure that the two aren’t conflated, but #Unicode 18.0 is still not rimshot-complete. We can start it off with the snare drum 🥁 but there will be no punchline to this setup. Are we, in the year 2026, to finish our emoji rimshots with a bell?? 🛎️

    The people deserve better.

    #WorldEmojiDay