home.social

#parsing — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #parsing, aggregated by home.social.

fetched live
  1. I know that PEG parsing is technically not linear-time like LALR(1), unless you use packrat parsing/memorization, but it's amazing how powerful such a simple, easily implemented parsing technique is. It is far simpler to implement than LALR(1) or LR(1), and for real-world use, the performance is perfectly acceptable.
    Unlike LALR(1), it has not been proven that there are CFGs that can't be parsed by PEG. In other words, there is no known CFG that cannot be parsed by a PEG.
    #Parsing

  2. I know that PEG parsing is technically not linear-time like LALR(1), unless you use packrat parsing/memorization, but it's amazing how powerful such a simple, easily implemented parsing technique is. It is far simpler to implement than LALR(1) or LR(1), and for real-world use, the performance is perfectly acceptable.
    Unlike LALR(1), it has not been proven that there are CFGs that can't be parsed by PEG. In other words, there is no known CFG that cannot be parsed by a PEG.
    #Parsing

  3. I know that PEG parsing is technically not linear-time like LALR(1), unless you use packrat parsing/memorization, but it's amazing how powerful such a simple, easily implemented parsing technique is. It is far simpler to implement than LALR(1) or LR(1), and for real-world use, the performance is perfectly acceptable.
    Unlike LALR(1), it has not been proven that there are CFGs that can't be parsed by PEG. In other words, there is no known CFG that cannot be parsed by a PEG.
    #Parsing

  4. I know that PEG parsing is technically not linear-time like LALR(1), unless you use packrat parsing/memorization, but it's amazing how powerful such a simple, easily implemented parsing technique is. It is far simpler to implement than LALR(1) or LR(1), and for real-world use, the performance is perfectly acceptable.
    Unlike LALR(1), it has not been proven that there are CFGs that can't be parsed by PEG. In other words, there is no known CFG that cannot be parsed by a PEG.
    #Parsing

  5. I know that PEG parsing is technically not linear-time like LALR(1), unless you use packrat parsing/memorization, but it's amazing how powerful such a simple, easily implemented parsing technique is. It is far simpler to implement than LALR(1) or LR(1), and for real-world use, the performance is perfectly acceptable.
    Unlike LALR(1), it has not been proven that there are CFGs that can't be parsed by PEG. In other words, there is no known CFG that cannot be parsed by a PEG.
    #Parsing

  6. 🚀🎉 Breaking #news, folks! The City of #Munich has decided to make history by sponsoring the riveting field of #XML #parsing for a whopping six months! 🙄 Watch out world, because with this groundbreaking investment, #libexpat will surely redefine the universe of open-source boredom. 🌐💤
    blog.hartwork.org/posts/libexp #OpenSource #HackerNews #ngated

  7. 🚀🎉 Breaking #news, folks! The City of #Munich has decided to make history by sponsoring the riveting field of #XML #parsing for a whopping six months! 🙄 Watch out world, because with this groundbreaking investment, #libexpat will surely redefine the universe of open-source boredom. 🌐💤
    blog.hartwork.org/posts/libexp #OpenSource #HackerNews #ngated

  8. 🚀🎉 Breaking #news, folks! The City of #Munich has decided to make history by sponsoring the riveting field of #XML #parsing for a whopping six months! 🙄 Watch out world, because with this groundbreaking investment, #libexpat will surely redefine the universe of open-source boredom. 🌐💤
    blog.hartwork.org/posts/libexp #OpenSource #HackerNews #ngated

  9. 🚀🎉 Breaking #news, folks! The City of #Munich has decided to make history by sponsoring the riveting field of #XML #parsing for a whopping six months! 🙄 Watch out world, because with this groundbreaking investment, #libexpat will surely redefine the universe of open-source boredom. 🌐💤
    blog.hartwork.org/posts/libexp #OpenSource #HackerNews #ngated

  10. 🚀🎉 Breaking #news, folks! The City of #Munich has decided to make history by sponsoring the riveting field of #XML #parsing for a whopping six months! 🙄 Watch out world, because with this groundbreaking investment, #libexpat will surely redefine the universe of open-source boredom. 🌐💤
    blog.hartwork.org/posts/libexp #OpenSource #HackerNews #ngated

  11. Language is overrated. In this Substack piece, I scrutinise language and the fact that much of language is extraneous. As for spoken language, many utterances are not strictly necessary, though they are functional.

    brywillis634737.substack.com/p

    I shake hands with #Austin, #Derrida, #Searle, #Wittgenstein, #Nietzsche, and #Barthes on this journey – even a bout with #Pinker.

    #philosophy #language #languageinsufficiency #salt #justice #blog #podcast #artificialintelligence #robots #parsing

  12. Language is overrated. In this Substack piece, I scrutinise language and the fact that much of language is extraneous. As for spoken language, many utterances are not strictly necessary, though they are functional.

    brywillis634737.substack.com/p

    I shake hands with #Austin, #Derrida, #Searle, #Wittgenstein, #Nietzsche, and #Barthes on this journey – even a bout with #Pinker.

    #philosophy #language #languageinsufficiency #salt #justice #blog #podcast #artificialintelligence #robots #parsing

  13. Language is overrated. In this Substack piece, I scrutinise language and the fact that much of language is extraneous. As for spoken language, many utterances are not strictly necessary, though they are functional.

    brywillis634737.substack.com/p

    I shake hands with #Austin, #Derrida, #Searle, #Wittgenstein, #Nietzsche, and #Barthes on this journey – even a bout with #Pinker.

    #philosophy #language #languageinsufficiency #salt #justice #blog #podcast #artificialintelligence #robots #parsing

  14. Language is overrated. In this Substack piece, I scrutinise language and the fact that much of language is extraneous. As for spoken language, many utterances are not strictly necessary, though they are functional.

    brywillis634737.substack.com/p

    I shake hands with #Austin, #Derrida, #Searle, #Wittgenstein, #Nietzsche, and #Barthes on this journey – even a bout with #Pinker.

    #philosophy #language #languageinsufficiency #salt #justice #blog #podcast #artificialintelligence #robots #parsing

  15. Language is overrated. In this Substack piece, I scrutinise language and the fact that much of language is extraneous. As for spoken language, many utterances are not strictly necessary, though they are functional.

    brywillis634737.substack.com/p

    I shake hands with #Austin, #Derrida, #Searle, #Wittgenstein, #Nietzsche, and #Barthes on this journey – even a bout with #Pinker.

    #philosophy #language #languageinsufficiency #salt #justice #blog #podcast #artificialintelligence #robots #parsing

  16. 🚀 Wow, a 2012 screed on regular expressions! Because nobody knew about their "true power" before this 🙄. Congratulations on discovering that #regex can sometimes parse HTML—surely a revelation for the ages. 🤦‍♂️ #RegexRevolution
    npopov.com/2012/06/15/The-true #RegularExpressions #HTML #Parsing #DeveloperHumor #TechNews #HackerNews #ngated

  17. 🚀 Wow, a 2012 screed on regular expressions! Because nobody knew about their "true power" before this 🙄. Congratulations on discovering that #regex can sometimes parse HTML—surely a revelation for the ages. 🤦‍♂️ #RegexRevolution
    npopov.com/2012/06/15/The-true #RegularExpressions #HTML #Parsing #DeveloperHumor #TechNews #HackerNews #ngated

  18. 🚀 Wow, a 2012 screed on regular expressions! Because nobody knew about their "true power" before this 🙄. Congratulations on discovering that #regex can sometimes parse HTML—surely a revelation for the ages. 🤦‍♂️ #RegexRevolution
    npopov.com/2012/06/15/The-true #RegularExpressions #HTML #Parsing #DeveloperHumor #TechNews #HackerNews #ngated

  19. 🚀 Wow, a 2012 screed on regular expressions! Because nobody knew about their "true power" before this 🙄. Congratulations on discovering that #regex can sometimes parse HTML—surely a revelation for the ages. 🤦‍♂️ #RegexRevolution
    npopov.com/2012/06/15/The-true #RegularExpressions #HTML #Parsing #DeveloperHumor #TechNews #HackerNews #ngated

  20. 🚀 Wow, a 2012 screed on regular expressions! Because nobody knew about their "true power" before this 🙄. Congratulations on discovering that #regex can sometimes parse HTML—surely a revelation for the ages. 🤦‍♂️ #RegexRevolution
    npopov.com/2012/06/15/The-true #RegularExpressions #HTML #Parsing #DeveloperHumor #TechNews #HackerNews #ngated

  21. Sigh... I guess I need to ask this now

    So #rust programmers, how does one do errors in lossless parsing?

    Context

    I'm working on the new nu parser for nushell. The current strategy is to have a giant vector of so called NodeId s which in turn refer other indices within that vector, and ultimately build a parse "tree" using those NodeId (variants in the AstNode may contain other metadata too)

    Now, we just created a dummy NodeId for errors and pushed them and continued to parse on our way. We want to have error-resilient parsing, so that we can cover more errors at other locations. If something's badly borked, well, we just barf.

    So this worked well for us devs, because it kept things simple easy, but this is not how I'd really like to build the parse tree. There are multiple problems.

    Right off the bat, we lose all the type information in one big ball of AstNode enum vector. So even though there are some places where you know that an AstNode would be of a particular kind, you simply cannot do anything about it except explicitly match and check (we want to avoid unsafe memory reinterprets).

    As an aside, most of the enum is empty, but due to some variants, the entire enum becomes 48 bytes, so we waste a lot of memory on basically nothing.

    My solution to this was to use bumpalo to do allocations and store references, and let go of the vector altogether and build an actual parse tree (like, a node with children, which have further children, and so on)

    This is basically doing the vector, but not wasting space on the emptiness of the enums, not lose the types at runtime to the enum, and not have to do hacky index-chasing (which effectively act like pointers now). Instead, we use actual references, which we know will be valid till the end of the arena lifetime (three cheers for bumpalo!)

    The Problem

    As mentioned before, we had NodeId for errors as well, and those had a dummy span. Now, with the existence of a concrete parse tree, we have... no such thing. So we lose out on the ease of creating an error and attaching it wherever we want. In order to do things the earlier way, we have to make everything into an Result<insert_type_here, Error> which is very, very bad, because with that comes the indirect habit of using ? to bubble upwards, which is a problem because we want to find as big of a parsed expression with as small of an error context (we basically want to minimize how much we weren't able to parse, because that's what any good parser should do)

    Right now, I don't really have any concrete ideas to attach errors to the nodes in the parse tree. So finally

    Questions

    • How does one do it? I'm looking for suggestions, although I don't really have anything concrete in my mind
    • Would this actually lead to better performance? I expect it to be both faster and cheaper resource-wise. I'd argue yes, but I have no idea since it's a massive rewrite and doesn't even compile yet (:

    Code snippets

    // NOTE(bumpalo_rewrite): This becomes the root of the AST
    // TODO(bumpalo_rewrite): Change to the generic enum that block may contain
    #[derive(Debug, Clone)]
    pub struct Block<'a> {
        pub span_start: usize,
        pub span_end: usize,
        pub nodes: Vec<BlockEntities<'a>>,
    }
    
    // TODO(bumpalo_rewrite): Fill with all the possible BlockEntities
    // NOTE(bumpalo_rewrite): See Parser::block() for the entity types
    #[derive(Debug, Clone)]
    pub enum BlockEntities<'a> {
        Def(&'a Def<'a>),
        Let(&'a Let<'a>),
        While(&'a While<'a>),
        For(&'a For<'a>),
        Loop(&'a Loop<'a>),
        Return(&'a Return<'a>),
        Continue(&'a Continue),
        Break(&'a Break),
        Alias(&'a Alias),
        Extern(&'a Extern),
        PipelineOrExprOrAssign(PipelineOrExprOrAssign<'a>),
        Statement(PipelineOrExprOrAssign<'a>),
    }
    

    This is kind of what it is like right now. The function that parses a block:

                else if self.is_keyword(b"while") {
                    match self.while_statement(arena) {
                        Some(while_) => {
                            code_body.push(BlockEntities::While(while_));
                        }
                        None => {}
                    };
                }
    

    has cases like these, where the error context is lost (ignore using Option instead of Result, I'm prototyping null

    So really, how do I preserve the error contexts?

    Git Repo

    Here. You'd want to mostly look through the diff between this commit and the previous one in src/parser.rs

    Any help is appreciated because I'm close to losing my mind lol XD

    #programming #parsing #rust #errortolerance #compilers #parsers

  22. Sigh... I guess I need to ask this now

    So #rust programmers, how does one do errors in lossless parsing?

    Context

    I'm working on the new nu parser for nushell. The current strategy is to have a giant vector of so called NodeId s which in turn refer other indices within that vector, and ultimately build a parse "tree" using those NodeId (variants in the AstNode may contain other metadata too)

    Now, we just created a dummy NodeId for errors and pushed them and continued to parse on our way. We want to have error-resilient parsing, so that we can cover more errors at other locations. If something's badly borked, well, we just barf.

    So this worked well for us devs, because it kept things simple easy, but this is not how I'd really like to build the parse tree. There are multiple problems.

    Right off the bat, we lose all the type information in one big ball of AstNode enum vector. So even though there are some places where you know that an AstNode would be of a particular kind, you simply cannot do anything about it except explicitly match and check (we want to avoid unsafe memory reinterprets).

    As an aside, most of the enum is empty, but due to some variants, the entire enum becomes 48 bytes, so we waste a lot of memory on basically nothing.

    My solution to this was to use bumpalo to do allocations and store references, and let go of the vector altogether and build an actual parse tree (like, a node with children, which have further children, and so on)

    This is basically doing the vector, but not wasting space on the emptiness of the enums, not lose the types at runtime to the enum, and not have to do hacky index-chasing (which effectively act like pointers now). Instead, we use actual references, which we know will be valid till the end of the arena lifetime (three cheers for bumpalo!)

    The Problem

    As mentioned before, we had NodeId for errors as well, and those had a dummy span. Now, with the existence of a concrete parse tree, we have... no such thing. So we lose out on the ease of creating an error and attaching it wherever we want. In order to do things the earlier way, we have to make everything into an Result<insert_type_here, Error> which is very, very bad, because with that comes the indirect habit of using ? to bubble upwards, which is a problem because we want to find as big of a parsed expression with as small of an error context (we basically want to minimize how much we weren't able to parse, because that's what any good parser should do)

    Right now, I don't really have any concrete ideas to attach errors to the nodes in the parse tree. So finally

    Questions

    • How does one do it? I'm looking for suggestions, although I don't really have anything concrete in my mind
    • Would this actually lead to better performance? I expect it to be both faster and cheaper resource-wise. I'd argue yes, but I have no idea since it's a massive rewrite and doesn't even compile yet (:

    Code snippets

    // NOTE(bumpalo_rewrite): This becomes the root of the AST
    // TODO(bumpalo_rewrite): Change to the generic enum that block may contain
    #[derive(Debug, Clone)]
    pub struct Block<'a> {
        pub span_start: usize,
        pub span_end: usize,
        pub nodes: Vec<BlockEntities<'a>>,
    }
    
    // TODO(bumpalo_rewrite): Fill with all the possible BlockEntities
    // NOTE(bumpalo_rewrite): See Parser::block() for the entity types
    #[derive(Debug, Clone)]
    pub enum BlockEntities<'a> {
        Def(&'a Def<'a>),
        Let(&'a Let<'a>),
        While(&'a While<'a>),
        For(&'a For<'a>),
        Loop(&'a Loop<'a>),
        Return(&'a Return<'a>),
        Continue(&'a Continue),
        Break(&'a Break),
        Alias(&'a Alias),
        Extern(&'a Extern),
        PipelineOrExprOrAssign(PipelineOrExprOrAssign<'a>),
        Statement(PipelineOrExprOrAssign<'a>),
    }
    

    This is kind of what it is like right now. The function that parses a block:

                else if self.is_keyword(b"while") {
                    match self.while_statement(arena) {
                        Some(while_) => {
                            code_body.push(BlockEntities::While(while_));
                        }
                        None => {}
                    };
                }
    

    has cases like these, where the error context is lost (ignore using Option instead of Result, I'm prototyping null

    So really, how do I preserve the error contexts?

    Git Repo

    Here. You'd want to mostly look through the diff between this commit and the previous one in src/parser.rs

    Any help is appreciated because I'm close to losing my mind lol XD

    #programming #parsing #rust #errortolerance #compilers #parsers

  23. Sigh... I guess I need to ask this now

    So #rust programmers, how does one do errors in lossless parsing?

    Context

    I'm working on the new nu parser for nushell. The current strategy is to have a giant vector of so called NodeId s which in turn refer other indices within that vector, and ultimately build a parse "tree" using those NodeId (variants in the AstNode may contain other metadata too)

    Now, we just created a dummy NodeId for errors and pushed them and continued to parse on our way. We want to have error-resilient parsing, so that we can cover more errors at other locations. If something's badly borked, well, we just barf.

    So this worked well for us devs, because it kept things simple easy, but this is not how I'd really like to build the parse tree. There are multiple problems.

    Right off the bat, we lose all the type information in one big ball of AstNode enum vector. So even though there are some places where you know that an AstNode would be of a particular kind, you simply cannot do anything about it except explicitly match and check (we want to avoid unsafe memory reinterprets).

    As an aside, most of the enum is empty, but due to some variants, the entire enum becomes 48 bytes, so we waste a lot of memory on basically nothing.

    My solution to this was to use bumpalo to do allocations and store references, and let go of the vector altogether and build an actual parse tree (like, a node with children, which have further children, and so on)

    This is basically doing the vector, but not wasting space on the emptiness of the enums, not lose the types at runtime to the enum, and not have to do hacky index-chasing (which effectively act like pointers now). Instead, we use actual references, which we know will be valid till the end of the arena lifetime (three cheers for bumpalo!)

    The Problem

    As mentioned before, we had NodeId for errors as well, and those had a dummy span. Now, with the existence of a concrete parse tree, we have... no such thing. So we lose out on the ease of creating an error and attaching it wherever we want. In order to do things the earlier way, we have to make everything into an Result<insert_type_here, Error> which is very, very bad, because with that comes the indirect habit of using ? to bubble upwards, which is a problem because we want to find as big of a parsed expression with as small of an error context (we basically want to minimize how much we weren't able to parse, because that's what any good parser should do)

    Right now, I don't really have any concrete ideas to attach errors to the nodes in the parse tree. So finally

    Questions

    • How does one do it? I'm looking for suggestions, although I don't really have anything concrete in my mind
    • Would this actually lead to better performance? I expect it to be both faster and cheaper resource-wise. I'd argue yes, but I have no idea since it's a massive rewrite and doesn't even compile yet (:

    Code snippets

    // NOTE(bumpalo_rewrite): This becomes the root of the AST
    // TODO(bumpalo_rewrite): Change to the generic enum that block may contain
    #[derive(Debug, Clone)]
    pub struct Block<'a> {
        pub span_start: usize,
        pub span_end: usize,
        pub nodes: Vec<BlockEntities<'a>>,
    }
    
    // TODO(bumpalo_rewrite): Fill with all the possible BlockEntities
    // NOTE(bumpalo_rewrite): See Parser::block() for the entity types
    #[derive(Debug, Clone)]
    pub enum BlockEntities<'a> {
        Def(&'a Def<'a>),
        Let(&'a Let<'a>),
        While(&'a While<'a>),
        For(&'a For<'a>),
        Loop(&'a Loop<'a>),
        Return(&'a Return<'a>),
        Continue(&'a Continue),
        Break(&'a Break),
        Alias(&'a Alias),
        Extern(&'a Extern),
        PipelineOrExprOrAssign(PipelineOrExprOrAssign<'a>),
        Statement(PipelineOrExprOrAssign<'a>),
    }
    

    This is kind of what it is like right now. The function that parses a block:

                else if self.is_keyword(b"while") {
                    match self.while_statement(arena) {
                        Some(while_) => {
                            code_body.push(BlockEntities::While(while_));
                        }
                        None => {}
                    };
                }
    

    has cases like these, where the error context is lost (ignore using Option instead of Result, I'm prototyping null

    So really, how do I preserve the error contexts?

    Git Repo

    Here. You'd want to mostly look through the diff between this commit and the previous one in src/parser.rs

    Any help is appreciated because I'm close to losing my mind lol XD

    #programming #parsing #rust #errortolerance #compilers #parsers

  24. Sigh... I guess I need to ask this now

    So #rust programmers, how does one do errors in lossless parsing?

    Context

    I'm working on the new nu parser for nushell. The current strategy is to have a giant vector of so called NodeId s which in turn refer other indices within that vector, and ultimately build a parse "tree" using those NodeId (variants in the AstNode may contain other metadata too)

    Now, we just created a dummy NodeId for errors and pushed them and continued to parse on our way. We want to have error-resilient parsing, so that we can cover more errors at other locations. If something's badly borked, well, we just barf.

    So this worked well for us devs, because it kept things simple easy, but this is not how I'd really like to build the parse tree. There are multiple problems.

    Right off the bat, we lose all the type information in one big ball of AstNode enum vector. So even though there are some places where you know that an AstNode would be of a particular kind, you simply cannot do anything about it except explicitly match and check (we want to avoid unsafe memory reinterprets).

    As an aside, most of the enum is empty, but due to some variants, the entire enum becomes 48 bytes, so we waste a lot of memory on basically nothing.

    My solution to this was to use bumpalo to do allocations and store references, and let go of the vector altogether and build an actual parse tree (like, a node with children, which have further children, and so on)

    This is basically doing the vector, but not wasting space on the emptiness of the enums, not lose the types at runtime to the enum, and not have to do hacky index-chasing (which effectively act like pointers now). Instead, we use actual references, which we know will be valid till the end of the arena lifetime (three cheers for bumpalo!)

    The Problem

    As mentioned before, we had NodeId for errors as well, and those had a dummy span. Now, with the existence of a concrete parse tree, we have... no such thing. So we lose out on the ease of creating an error and attaching it wherever we want. In order to do things the earlier way, we have to make everything into an Result<insert_type_here, Error> which is very, very bad, because with that comes the indirect habit of using ? to bubble upwards, which is a problem because we want to find as big of a parsed expression with as small of an error context (we basically want to minimize how much we weren't able to parse, because that's what any good parser should do)

    Right now, I don't really have any concrete ideas to attach errors to the nodes in the parse tree. So finally

    Questions

    • How does one do it? I'm looking for suggestions, although I don't really have anything concrete in my mind
    • Would this actually lead to better performance? I expect it to be both faster and cheaper resource-wise. I'd argue yes, but I have no idea since it's a massive rewrite and doesn't even compile yet (:

    Code snippets

    // NOTE(bumpalo_rewrite): This becomes the root of the AST
    // TODO(bumpalo_rewrite): Change to the generic enum that block may contain
    #[derive(Debug, Clone)]
    pub struct Block<'a> {
        pub span_start: usize,
        pub span_end: usize,
        pub nodes: Vec<BlockEntities<'a>>,
    }
    
    // TODO(bumpalo_rewrite): Fill with all the possible BlockEntities
    // NOTE(bumpalo_rewrite): See Parser::block() for the entity types
    #[derive(Debug, Clone)]
    pub enum BlockEntities<'a> {
        Def(&'a Def<'a>),
        Let(&'a Let<'a>),
        While(&'a While<'a>),
        For(&'a For<'a>),
        Loop(&'a Loop<'a>),
        Return(&'a Return<'a>),
        Continue(&'a Continue),
        Break(&'a Break),
        Alias(&'a Alias),
        Extern(&'a Extern),
        PipelineOrExprOrAssign(PipelineOrExprOrAssign<'a>),
        Statement(PipelineOrExprOrAssign<'a>),
    }
    

    This is kind of what it is like right now. The function that parses a block:

                else if self.is_keyword(b"while") {
                    match self.while_statement(arena) {
                        Some(while_) => {
                            code_body.push(BlockEntities::While(while_));
                        }
                        None => {}
                    };
                }
    

    has cases like these, where the error context is lost (ignore using Option instead of Result, I'm prototyping null

    So really, how do I preserve the error contexts?

    Git Repo

    Here. You'd want to mostly look through the diff between this commit and the previous one in src/parser.rs

    Any help is appreciated because I'm close to losing my mind lol XD

    #programming #parsing #rust #errortolerance #compilers #parsers

  25. PR pour rendre meilleure la prise en charge du kabyle sur ce « number parser » d'Open Voice OS.

    Faut savoir que même `libnumbertext` ne prend pas en charge le système de numération en kabyle. J'ai une version localement qui le supporte.

    En attendant :

    github.com/OpenVoiceOS/ovos-nu

    #kabyle #numbers #numerals #parsing

  26. PR pour rendre meilleure la prise en charge du kabyle sur ce « number parser » d'Open Voice OS.

    Faut savoir que même `libnumbertext` ne prend pas en charge le système de numération en kabyle. J'ai une version localement qui le supporte.

    En attendant :

    github.com/OpenVoiceOS/ovos-nu

    #kabyle #numbers #numerals #parsing

  27. PR pour rendre meilleure la prise en charge du kabyle sur ce « number parser » d'Open Voice OS.

    Faut savoir que même `libnumbertext` ne prend pas en charge le système de numération en kabyle. J'ai une version localement qui le supporte.

    En attendant :

    github.com/OpenVoiceOS/ovos-nu

    #kabyle #numbers #numerals #parsing

  28. PR pour rendre meilleure la prise en charge du kabyle sur ce « number parser » d'Open Voice OS.

    Faut savoir que même `libnumbertext` ne prend pas en charge le système de numération en kabyle. J'ai une version localement qui le supporte.

    En attendant :

    github.com/OpenVoiceOS/ovos-nu

    #kabyle #numbers #numerals #parsing

  29. PR pour rendre meilleure la prise en charge du kabyle sur ce « number parser » d'Open Voice OS.

    Faut savoir que même `libnumbertext` ne prend pas en charge le système de numération en kabyle. J'ai une version localement qui le supporte.

    En attendant :

    github.com/OpenVoiceOS/ovos-nu

    #kabyle #numbers #numerals #parsing

  30. Как я случайно написал что-то быстрое и декларативное (на Rust)

    Писал парсер строго под свой проект, а получился быстрый декларативный движок для парсинга текстовых форматов. Как?

    habr.com/ru/articles/1049432/

    #rust #parsing #engine #declarative #fast #парсер #парсинг #работа_с_текстом #работа_с_текстовыми_данными

  31. Как я случайно написал что-то быстрое и декларативное (на Rust)

    Писал парсер строго под свой проект, а получился быстрый декларативный движок для парсинга текстовых форматов. Как?

    habr.com/ru/articles/1049432/

    #rust #parsing #engine #declarative #fast #парсер #парсинг #работа_с_текстом #работа_с_текстовыми_данными

  32. Как я случайно написал что-то быстрое и декларативное (на Rust)

    Писал парсер строго под свой проект, а получился быстрый декларативный движок для парсинга текстовых форматов. Как?

    habr.com/ru/articles/1049432/

    #rust #parsing #engine #declarative #fast #парсер #парсинг #работа_с_текстом #работа_с_текстовыми_данными

  33. Okay, es ist mal wieder Psycholinguistik-Time, weil #Psycholinguistik geilste #Linguistik!

    Lest mal den folgenden Satz:

    "Dass Nagelsmann zugunsten von Sané nie etwas unternommen worden wäre, ist eine Lüge."

    Ich gebe zu, der ist ein bisschen kompliziert, und vielleicht habt Ihr gedacht: "Müsste da statt 'worden wäre' nicht 'hat' stehen?" - Ja, kann es, aber es ist auch so ein auflösbarer ("richtiger") Satz.

    Interessant ist aber, dass es in der Form oben ein #Holzweg-Satz (Engl. #gardenpath) ist. Die heißen deshalb so, weil sie einen bei der Satzinterpretation (Engl. #parsing) in die Irre führen: Die meisten Leute denken beim ersten Lesen, dass Nagelsmann die handelnde Person (Agens) beim "unternehmen" ist. In Wahrheit ist es aber Sané – und das wird klar, wenn man den Satz entsprechend betont. Liegt die Betonung auf Nagelsmann, wird die Interpretation einfacher.

  34. Okay, es ist mal wieder Psycholinguistik-Time, weil #Psycholinguistik geilste #Linguistik!

    Lest mal den folgenden Satz:

    "Dass Nagelsmann zugunsten von Sané nie etwas unternommen worden wäre, ist eine Lüge."

    Ich gebe zu, der ist ein bisschen kompliziert, und vielleicht habt Ihr gedacht: "Müsste da statt 'worden wäre' nicht 'hat' stehen?" - Ja, kann es, aber es ist auch so ein auflösbarer ("richtiger") Satz.

    Interessant ist aber, dass es in der Form oben ein #Holzweg-Satz (Engl. #gardenpath) ist. Die heißen deshalb so, weil sie einen bei der Satzinterpretation (Engl. #parsing) in die Irre führen: Die meisten Leute denken beim ersten Lesen, dass Nagelsmann die handelnde Person (Agens) beim "unternehmen" ist. In Wahrheit ist es aber Sané – und das wird klar, wenn man den Satz entsprechend betont. Liegt die Betonung auf Nagelsmann, wird die Interpretation einfacher.

  35. Okay, es ist mal wieder Psycholinguistik-Time, weil #Psycholinguistik geilste #Linguistik!

    Lest mal den folgenden Satz:

    "Dass Nagelsmann zugunsten von Sané nie etwas unternommen worden wäre, ist eine Lüge."

    Ich gebe zu, der ist ein bisschen kompliziert, und vielleicht habt Ihr gedacht: "Müsste da statt 'worden wäre' nicht 'hat' stehen?" - Ja, kann es, aber es ist auch so ein auflösbarer ("richtiger") Satz.

    Interessant ist aber, dass es in der Form oben ein #Holzweg-Satz (Engl. #gardenpath) ist. Die heißen deshalb so, weil sie einen bei der Satzinterpretation (Engl. #parsing) in die Irre führen: Die meisten Leute denken beim ersten Lesen, dass Nagelsmann die handelnde Person (Agens) beim "unternehmen" ist. In Wahrheit ist es aber Sané – und das wird klar, wenn man den Satz entsprechend betont. Liegt die Betonung auf Nagelsmann, wird die Interpretation einfacher.

  36. Okay, es ist mal wieder Psycholinguistik-Time, weil #Psycholinguistik geilste #Linguistik!

    Lest mal den folgenden Satz:

    "Dass Nagelsmann zugunsten von Sané nie etwas unternommen worden wäre, ist eine Lüge."

    Ich gebe zu, der ist ein bisschen kompliziert, und vielleicht habt Ihr gedacht: "Müsste da statt 'worden wäre' nicht 'hat' stehen?" - Ja, kann es, aber es ist auch so ein auflösbarer ("richtiger") Satz.

    Interessant ist aber, dass es in der Form oben ein #Holzweg-Satz (Engl. #gardenpath) ist. Die heißen deshalb so, weil sie einen bei der Satzinterpretation (Engl. #parsing) in die Irre führen: Die meisten Leute denken beim ersten Lesen, dass Nagelsmann die handelnde Person (Agens) beim "unternehmen" ist. In Wahrheit ist es aber Sané – und das wird klar, wenn man den Satz entsprechend betont. Liegt die Betonung auf Nagelsmann, wird die Interpretation einfacher.

  37. Okay, es ist mal wieder Psycholinguistik-Time, weil #Psycholinguistik geilste #Linguistik!

    Lest mal den folgenden Satz:

    "Dass Nagelsmann zugunsten von Sané nie etwas unternommen worden wäre, ist eine Lüge."

    Ich gebe zu, der ist ein bisschen kompliziert, und vielleicht habt Ihr gedacht: "Müsste da statt 'worden wäre' nicht 'hat' stehen?" - Ja, kann es, aber es ist auch so ein auflösbarer ("richtiger") Satz.

    Interessant ist aber, dass es in der Form oben ein #Holzweg-Satz (Engl. #gardenpath) ist. Die heißen deshalb so, weil sie einen bei der Satzinterpretation (Engl. #parsing) in die Irre führen: Die meisten Leute denken beim ersten Lesen, dass Nagelsmann die handelnde Person (Agens) beim "unternehmen" ist. In Wahrheit ist es aber Sané – und das wird klar, wenn man den Satz entsprechend betont. Liegt die Betonung auf Nagelsmann, wird die Interpretation einfacher.

  38. Шифрование на уровне протокола

    Как организовать шифрование на уровне протокола? На самом деле тема непростая и пожалуй (имхо) это как раз та самая тема, где прийти к компромиссу почти никогда не получается. Разве что просто не передавать чувствительные данные вовсе. Я расскажу как шифрование можно организовать на уровне протокола brec и ни в коем случае не буду затрагивать те самые принципиальные решения, влияющие на безопасность (как передавать, куда передавать, отправлять ли, и хранить ли чувствительные данные вовсе). Иными словами нас интересует инструментальная сторона вопроса.

    habr.com/ru/articles/1040298/

    #rust #protocols #parsing #binary #communication #no_database #storage #binary_storage #filtering #crypt

  39. Шифрование на уровне протокола

    Как организовать шифрование на уровне протокола? На самом деле тема непростая и пожалуй (имхо) это как раз та самая тема, где прийти к компромиссу почти никогда не получается. Разве что просто не передавать чувствительные данные вовсе. Я расскажу как шифрование можно организовать на уровне протокола brec и ни в коем случае не буду затрагивать те самые принципиальные решения, влияющие на безопасность (как передавать, куда передавать, отправлять ли, и хранить ли чувствительные данные вовсе). Иными словами нас интересует инструментальная сторона вопроса.

    habr.com/ru/articles/1040298/

    #rust #protocols #parsing #binary #communication #no_database #storage #binary_storage #filtering #crypt

  40. Шифрование на уровне протокола

    Как организовать шифрование на уровне протокола? На самом деле тема непростая и пожалуй (имхо) это как раз та самая тема, где прийти к компромиссу почти никогда не получается. Разве что просто не передавать чувствительные данные вовсе. Я расскажу как шифрование можно организовать на уровне протокола brec и ни в коем случае не буду затрагивать те самые принципиальные решения, влияющие на безопасность (как передавать, куда передавать, отправлять ли, и хранить ли чувствительные данные вовсе). Иными словами нас интересует инструментальная сторона вопроса.

    habr.com/ru/articles/1040298/

    #rust #protocols #parsing #binary #communication #no_database #storage #binary_storage #filtering #crypt

  41. 🎉 Behold the ultimate #toolkit for #nerds who find #parsing MPEG-TS streams exhilarating! #TSDuck presents a dizzying array of standards and protocols, because nothing screams #fun like endless acronyms and a PDF bonanza 📚. Dive in, if you dare, and unleash your inner transport stream aficionado! 😂
    tsduck.io/ #MPEGTS #TransportStream #HackerNews #ngated

  42. 🎉 Behold the ultimate #toolkit for #nerds who find #parsing MPEG-TS streams exhilarating! #TSDuck presents a dizzying array of standards and protocols, because nothing screams #fun like endless acronyms and a PDF bonanza 📚. Dive in, if you dare, and unleash your inner transport stream aficionado! 😂
    tsduck.io/ #MPEGTS #TransportStream #HackerNews #ngated

  43. 🎉 Behold the ultimate #toolkit for #nerds who find #parsing MPEG-TS streams exhilarating! #TSDuck presents a dizzying array of standards and protocols, because nothing screams #fun like endless acronyms and a PDF bonanza 📚. Dive in, if you dare, and unleash your inner transport stream aficionado! 😂
    tsduck.io/ #MPEGTS #TransportStream #HackerNews #ngated

  44. 🎉 Behold the ultimate #toolkit for #nerds who find #parsing MPEG-TS streams exhilarating! #TSDuck presents a dizzying array of standards and protocols, because nothing screams #fun like endless acronyms and a PDF bonanza 📚. Dive in, if you dare, and unleash your inner transport stream aficionado! 😂
    tsduck.io/ #MPEGTS #TransportStream #HackerNews #ngated

  45. 🎉 Behold the ultimate #toolkit for #nerds who find #parsing MPEG-TS streams exhilarating! #TSDuck presents a dizzying array of standards and protocols, because nothing screams #fun like endless acronyms and a PDF bonanza 📚. Dive in, if you dare, and unleash your inner transport stream aficionado! 😂
    tsduck.io/ #MPEGTS #TransportStream #HackerNews #ngated

  46. The fastest way to match characters on ARM processors?, lemire.me/blog/2026/04/19/the-.

    In this article, Lemir talks about two SIMD ARM SVE/SVE2 instructions: `match` and `nmatch`, which fit nicely in the _vectorized classification_ step of `simdjson`. These instructions improve the performance of `simdjson` from 11.4Gb/s to 14.4Gb/s.

    #performance #simd #arm #json #parsing