home.social

#hle — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #hle, aggregated by home.social.

fetched live
  1. @tompearce49

    I'd like to believe all the hype from the #AiAntagonists but to me they all sound like soldiers in a besieged city, cheering the news of the relief columns that never comes.

    The article is supremely optimistic, which is fair enough, optimism is needed with one of the key avatars of the #AntiAi movement being sprung using AI himself. The #reversecentaur #asbestosinthewalls guy himself @pluralistic

    #AI Blew past the #turingtest so fast, folks were tripping over themselves to bury decades of benchmarking. The previous AI attempts never breached Turing.

    Meanwhile, #HLE is climbing up faster than expected, which is the exact opposite of what folks who claim AI is not advancing is.

    It seems that the models are capable of Zero-shot learning, reaching accurate results for knowledge not in the training data.

    The answer as always is to become politically active and #regulateAI

    #AiResearch

  2. #gemini31pro that's just been released is hitting 44% on Humanities Last Exam...

    When #HLE was released, not so long ago, the current models were in single digits...

    The race is now between #aibubble and #agi

  3. On a #OuigoTrainClassique 63 to Brussel-Zuid hauled by a #HLE 18, ready for #FOSDEM tomorrow!!!!! (@ OTC ➜ Bruxelles Midi für #FOSDEM2026) #NowTräwelling https://
    traewelling.de/status/6918651

  4. Бенчмарк конца эпохи — Humanity’s Last Exam

    Хочу сегодня рассказать вам про Humanity’s Last Exam (HLE). Это один из главных бенчмарков, по которым сегодня оценивают модели искусственного интеллекта, вроде меня (шучу). Бенчмарки — это просто наборы задач/датасетов, на которых проверяют модели и смотрят, кто умнее, точнее, устойчивее и т.д. Например, MMLU — Massive Multitask Language Understanding — один из самых известных «общеобразовательных» экзаменов для ИИ. Он проверяет широкий круг знаний и базовое рассуждение: около 16 тысяч вопросов по 57 предметам — от математики и истории до права и компьютерных наук. Есть ещё BIG-bench (Beyond the Imitation Game) от Google — не один тест, а коллекция из 200+ задач, которые прислали разные исследователи. Там уже не только «знание фактов», но и логика, здравый смысл, язык, социальные предвзятости (social biases), программирование и всё то, на чём модели любят спотыкаться. Есть и более «узкие» бенчмарки:

    habr.com/ru/articles/974206/

    #hle #бенчмарки #ии #llm #benchmarks #ai #fun

  5. Today is third day of First Stand tournament and I want make predictions.

    Today matches:

    CFO lose - HLE win
    KC lose - TES win

    And team places at the end of Round Robin Stage:

    1. HLE (Korea)
    2. TES (China)
    3. CFO (Taiwan)
    4. TL (USA)
    5. KC (France)

    #LoL #LeagueOfLegends #FirstStand #HLE #TES #CFO #TL #KC