home.social

#deepread — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #deepread, aggregated by home.social.

fetched live
  1. Back in the days of 2021
    there was a lovely evaluation paper:
    Automatically identifying label errors
    Improving score's reliability
    Finding example's difficulty
    Active Learning

    aclanthology.org/2021.acl-long

    @par @hoyle
    #machinelearning #evaluation #IRT #LLM #deepRead

  2. Back in the days of 2021
    there was a lovely evaluation paper:
    Automatically identifying label errors
    Improving score's reliability
    Finding example's difficulty
    Active Learning

    aclanthology.org/2021.acl-long

    @par @hoyle
    #machinelearning #evaluation #IRT #LLM #deepRead

  3. Back in the days of 2021
    there was a lovely evaluation paper:
    Automatically identifying label errors
    Improving score's reliability
    Finding example's difficulty
    Active Learning

    aclanthology.org/2021.acl-long

    @par @hoyle
    #machinelearning #evaluation #IRT #LLM #deepRead

  4. Back in the days of 2021
    there was a lovely evaluation paper:
    Automatically identifying label errors
    Improving score's reliability
    Finding example's difficulty
    Active Learning

    aclanthology.org/2021.acl-long

    @par @hoyle
    #machinelearning #evaluation #IRT #LLM #deepRead

  5. Back in the days of 2021
    there was a lovely evaluation paper:
    Automatically identifying label errors
    Improving score's reliability
    Finding example's difficulty
    Active Learning

    aclanthology.org/2021.acl-long

    @par @hoyle
    #machinelearning #evaluation #IRT #LLM #deepRead

  6. Did you know:
    Evaluating a single model on HELM took
    ⏱️4K GPU hours or 💸+10K$ in API calls?!
    Flash-HELM⚡️can reduce costs by X200!
    arxiv.org/abs/2308.11696

    #deepRead #machinelearning #evaluation #eval #nlproc #NLP #LLM

  7. Did you know:
    Evaluating a single model on HELM took
    ⏱️4K GPU hours or 💸+10K$ in API calls?!
    Flash-HELM⚡️can reduce costs by X200!
    arxiv.org/abs/2308.11696

    #deepRead #machinelearning #evaluation #eval #nlproc #NLP #LLM

  8. Did you know:
    Evaluating a single model on HELM took
    ⏱️4K GPU hours or 💸+10K$ in API calls?!
    Flash-HELM⚡️can reduce costs by X200!
    arxiv.org/abs/2308.11696

    #deepRead #machinelearning #evaluation #eval #nlproc #NLP #LLM

  9. Did you know:
    Evaluating a single model on HELM took
    ⏱️4K GPU hours or 💸+10K$ in API calls?!
    Flash-HELM⚡️can reduce costs by X200!
    arxiv.org/abs/2308.11696

    #deepRead #machinelearning #evaluation #eval #nlproc #NLP #LLM

  10. Did you know:
    Evaluating a single model on HELM took
    ⏱️4K GPU hours or 💸+10K$ in API calls?!
    Flash-HELM⚡️can reduce costs by X200!
    arxiv.org/abs/2308.11696

    #deepRead #machinelearning #evaluation #eval #nlproc #NLP #LLM

  11. The newFormer is introduced,
    but what do we really know about it?

    @ari and others
    imagine a new large-scale architecture &
    ask how would you interptret its abilities and behaviours 🧵
    arxiv.org/abs/2308.00189
    #deepRead #NLProc #MachineLearning

  12. The newFormer is introduced,
    but what do we really know about it?

    @ari and others
    imagine a new large-scale architecture &
    ask how would you interptret its abilities and behaviours 🧵
    arxiv.org/abs/2308.00189
    #deepRead #NLProc #MachineLearning

  13. The newFormer is introduced,
    but what do we really know about it?

    @ari and others
    imagine a new large-scale architecture &
    ask how would you interptret its abilities and behaviours 🧵
    arxiv.org/abs/2308.00189
    #deepRead #NLProc #MachineLearning

  14. The newFormer is introduced,
    but what do we really know about it?

    @ari and others
    imagine a new large-scale architecture &
    ask how would you interptret its abilities and behaviours 🧵
    arxiv.org/abs/2308.00189
    #deepRead #NLProc #MachineLearning

  15. The newFormer is introduced,
    but what do we really know about it?

    @ari and others
    imagine a new large-scale architecture &
    ask how would you interptret its abilities and behaviours 🧵
    arxiv.org/abs/2308.00189
    #deepRead #NLProc #MachineLearning

  16. @mega Linear transformations can skip over layers, even till the end

    We can see 👀 what the network 🧠 thought!
    We can stop🛑 generating at early layers!

    arxiv.org/abs/2303.09435v1

    #NLProc #deepRead

  17. @mega Linear transformations can skip over layers, even till the end

    We can see 👀 what the network 🧠 thought!
    We can stop🛑 generating at early layers!

    arxiv.org/abs/2303.09435v1

    #NLProc #deepRead

  18. @mega Linear transformations can skip over layers, even till the end

    We can see 👀 what the network 🧠 thought!
    We can stop🛑 generating at early layers!

    arxiv.org/abs/2303.09435v1

    #NLProc #deepRead

  19. @mega Linear transformations can skip over layers, even till the end

    We can see 👀 what the network 🧠 thought!
    We can stop🛑 generating at early layers!

    arxiv.org/abs/2303.09435v1

    #NLProc #deepRead

  20. @mega Linear transformations can skip over layers, even till the end

    We can see 👀 what the network 🧠 thought!
    We can stop🛑 generating at early layers!

    arxiv.org/abs/2303.09435v1

    #NLProc #deepRead

  21. 🔎What's in a layer?🌹🕵🏻‍♀️

    Representations are vectors
    If only they were words...

    Finding:
    Any layer can be mapped well to another linearly
    Simple, efficient & interpretable
    & improves early exit

    arxiv.org/abs/2303.09435v1
    Story and 🧵
    #nlproc #deepRead #MachinLearning

  22. 🔎What's in a layer?🌹🕵🏻‍♀️

    Representations are vectors
    If only they were words...

    Finding:
    Any layer can be mapped well to another linearly
    Simple, efficient & interpretable
    & improves early exit

    arxiv.org/abs/2303.09435v1
    Story and 🧵
    #nlproc #deepRead #MachinLearning

  23. 🔎What's in a layer?🌹🕵🏻‍♀️

    Representations are vectors
    If only they were words...

    Finding:
    Any layer can be mapped well to another linearly
    Simple, efficient & interpretable
    & improves early exit

    arxiv.org/abs/2303.09435v1
    Story and 🧵
    #nlproc #deepRead #MachinLearning

  24. 🔎What's in a layer?🌹🕵🏻‍♀️

    Representations are vectors
    If only they were words...

    Finding:
    Any layer can be mapped well to another linearly
    Simple, efficient & interpretable
    & improves early exit

    arxiv.org/abs/2303.09435v1
    Story and 🧵
    #nlproc #deepRead #MachinLearning

  25. 🔎What's in a layer?🌹🕵🏻‍♀️

    Representations are vectors
    If only they were words...

    Finding:
    Any layer can be mapped well to another linearly
    Simple, efficient & interpretable
    & improves early exit

    arxiv.org/abs/2303.09435v1
    Story and 🧵
    #nlproc #deepRead #MachinLearning

  26. Mindblowing pretraining paradigm

    Train the same model to predict the two directions separately
    Better results, more parallelization

    arxiv.org/abs/2303.07295
    #deepRead #nlproc #pretraining #machinelearning

  27. Mindblowing pretraining paradigm

    Train the same model to predict the two directions separately
    Better results, more parallelization

    arxiv.org/abs/2303.07295
    #deepRead #nlproc #pretraining #machinelearning

  28. Mindblowing pretraining paradigm

    Train the same model to predict the two directions separately
    Better results, more parallelization

    arxiv.org/abs/2303.07295
    #deepRead #nlproc #pretraining #machinelearning

  29. Mindblowing pretraining paradigm

    Train the same model to predict the two directions separately
    Better results, more parallelization

    arxiv.org/abs/2303.07295
    #deepRead #nlproc #pretraining #machinelearning

  30. Mindblowing pretraining paradigm

    Train the same model to predict the two directions separately
    Better results, more parallelization

    arxiv.org/abs/2303.07295
    #deepRead #nlproc #pretraining #machinelearning

  31. 3 reasons for hallucinations started
    only 2 prevailed

    Finding how networks behave while hallucinating, they
    filter hallucinations (with great success)

    arxiv.org/abs/2301.07779
    #NLProc #neuralEmpty #NLP #deepRead

  32. 3 reasons for hallucinations started
    only 2 prevailed

    Finding how networks behave while hallucinating, they
    filter hallucinations (with great success)

    arxiv.org/abs/2301.07779
    #NLProc #neuralEmpty #NLP #deepRead

  33. 3 reasons for hallucinations started
    only 2 prevailed

    Finding how networks behave while hallucinating, they
    filter hallucinations (with great success)

    arxiv.org/abs/2301.07779
    #NLProc #neuralEmpty #NLP #deepRead

  34. 3 reasons for hallucinations started
    only 2 prevailed

    Finding how networks behave while hallucinating, they
    filter hallucinations (with great success)

    arxiv.org/abs/2301.07779
    #NLProc #neuralEmpty #NLP #deepRead

  35. 3 reasons for hallucinations started
    only 2 prevailed

    Finding how networks behave while hallucinating, they
    filter hallucinations (with great success)

    arxiv.org/abs/2301.07779
    #NLProc #neuralEmpty #NLP #deepRead

  36. What neurons determine agreement in multilingual LLMs?

    #deepRead but some answers:
    Across languages-2 distinct ways to encode syntax
    Share neurons not info

    Autoregressive have dedicated synt. neurons (MLM just spread across)

    @[email protected] yu xia @[email protected] #conllLivetweet2022

  37. What neurons determine agreement in multilingual LLMs?

    #deepRead but some answers:
    Across languages-2 distinct ways to encode syntax
    Share neurons not info

    Autoregressive have dedicated synt. neurons (MLM just spread across)

    @[email protected] yu xia @[email protected] #conllLivetweet2022

  38. What neurons determine agreement in multilingual LLMs?

    #deepRead but some answers:
    Across languages-2 distinct ways to encode syntax
    Share neurons not info

    Autoregressive have dedicated synt. neurons (MLM just spread across)

    @[email protected] yu xia @[email protected] #conllLivetweet2022

  39. What neurons determine agreement in multilingual LLMs?

    #deepRead but some answers:
    Across languages-2 distinct ways to encode syntax
    Share neurons not info

    Autoregressive have dedicated synt. neurons (MLM just spread across)

    @[email protected] yu xia @[email protected] #conllLivetweet2022

  40. What neurons determine agreement in multilingual LLMs?

    #deepRead but some answers:
    Across languages-2 distinct ways to encode syntax
    Share neurons not info

    Autoregressive have dedicated synt. neurons (MLM just spread across)

    @[email protected] yu xia @[email protected] #conllLivetweet2022