home.social

#jev — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #jev, aggregated by home.social.

  1. #Cloudflare released #Clef and #Clefflash #decisionmodels, claiming faster, more accurate performance than #Jev via #Qwen backbones. Clef handles images and video, supports 64k context, costs $0.24 per million tokens, and is Jev-compatible on Workers AI or Hugging Face. theregister.com/ai-and-ml/2026 #tech #news #ainews

  2. #Cloudflare released #Clef and #Clefflash #decisionmodels, claiming faster, more accurate performance than #Jev via #Qwen backbones. Clef handles images and video, supports 64k context, costs $0.24 per million tokens, and is Jev-compatible on Workers AI or Hugging Face. theregister.com/ai-and-ml/2026 #tech #news #ainews

  3. #Cloudflare released #Clef and #Clefflash #decisionmodels, claiming faster, more accurate performance than #Jev via #Qwen backbones. Clef handles images and video, supports 64k context, costs $0.24 per million tokens, and is Jev-compatible on Workers AI or Hugging Face. theregister.com/ai-and-ml/2026 #tech #news #ainews

  4. #Cloudflare released #Clef and #Clefflash #decisionmodels, claiming faster, more accurate performance than #Jev via #Qwen backbones. Clef handles images and video, supports 64k context, costs $0.24 per million tokens, and is Jev-compatible on Workers AI or Hugging Face. theregister.com/ai-and-ml/2026 #tech #news #ainews

  5. New post: Laya vs Jev, ten days later 🧠⚖️

    Ten days ago I benchmarked two "System One" decision models for my LLM router: Jev (TypeSafe, hosted) was usable as-is, Laya (open weights) zero-shot was not. Laya's README says "fine-tune me", so I did, on a 16 GB M4 MacBook.

    🔁 Laya 0.3.6 → 0.3.24: 18 releases, but the weights are byte-identical. Same numbers.

    🆕 New open rivals on my 80-prompt bench:
    • Laya typed-decisions: 71% topic, but never confident → over-provisions everything
    • Von 1.3: closest to gold out of the box (48%), but over-confident
    • Kev-0.8B: 81% topic, 100% multilingual, very under-confident
    Jev 1.13: 89% topic, 72.5% same route as gold, still the only one usable as-is.

    🎯 Fine-tuning Laya on my own labels (not Jev's: distilling Jev is against TypeSafe's terms):
    • 480 labelled prompts → 1,920 training items
    • official Apple Silicon script, MPS
    • 41.5 min for 4 epochs, 15.4 GB peak memory
    • rule of thumb: ~6 s per labelled prompt

    Results:
    • topic 59% → 85%
    • complexity 45% → 72.5% (= Jev)
    • risk 31% → 69% (> Jev)
    • but 0% confident answers 🙃 → a temperature fitted on held-out data (T=0.87) fixes it
    • same route as gold: 20% → 64% (Jev 72.5%), routed cost $1.28 vs Jev's $1.35

    Caveats in the post: same labeller for train and test, tiny groups, and it under-provisions 9 prompts vs 3 for Jev.

    🏠 Then the reality check: I replayed 80 real Home Assistant decisions (VMC, shutters, water, watering) through Jev and stock Laya.
    • Laya English answers ~0.5 to everything, and cuts 64/80 requests (512-token window; my prompts grew to 1–1.9k tokens with history)
    • Laya multilingual says "yes" to almost everything, including "normal" (0.97) to a night-time leak 💧
    • Jev in the house: VMC disagreements with the rule went from 100% to 6% once it got history. Half a cent per day.

    So the house stays on Jev. The local path is clear though: fine-tune on my own feedback labels, not on Jev's answers.

    Full post, commands, outputs and timings (EN/FR/IT):
    mornati.net/laya-vs-jev-ten-da

    Code & recipe: github.com/mmornati/system-one

    Thanks again to everyone who commented on the last Home Assistant post: the "why not a small local model?" question is what pushed me to measure this properly.

    #AI #LLM #FineTuning #HomeAssistant #SmartHome #AppleSilicon #OpenSource #Jev #Laya #SelfHosted

  6. New post: Laya vs Jev, ten days later 🧠⚖️

    Ten days ago I benchmarked two "System One" decision models for my LLM router: Jev (TypeSafe, hosted) was usable as-is, Laya (open weights) zero-shot was not. Laya's README says "fine-tune me", so I did, on a 16 GB M4 MacBook.

    🔁 Laya 0.3.6 → 0.3.24: 18 releases, but the weights are byte-identical. Same numbers.

    🆕 New open rivals on my 80-prompt bench:
    • Laya typed-decisions: 71% topic, but never confident → over-provisions everything
    • Von 1.3: closest to gold out of the box (48%), but over-confident
    • Kev-0.8B: 81% topic, 100% multilingual, very under-confident
    Jev 1.13: 89% topic, 72.5% same route as gold, still the only one usable as-is.

    🎯 Fine-tuning Laya on my own labels (not Jev's: distilling Jev is against TypeSafe's terms):
    • 480 labelled prompts → 1,920 training items
    • official Apple Silicon script, MPS
    • 41.5 min for 4 epochs, 15.4 GB peak memory
    • rule of thumb: ~6 s per labelled prompt

    Results:
    • topic 59% → 85%
    • complexity 45% → 72.5% (= Jev)
    • risk 31% → 69% (> Jev)
    • but 0% confident answers 🙃 → a temperature fitted on held-out data (T=0.87) fixes it
    • same route as gold: 20% → 64% (Jev 72.5%), routed cost $1.28 vs Jev's $1.35

    Caveats in the post: same labeller for train and test, tiny groups, and it under-provisions 9 prompts vs 3 for Jev.

    🏠 Then the reality check: I replayed 80 real Home Assistant decisions (VMC, shutters, water, watering) through Jev and stock Laya.
    • Laya English answers ~0.5 to everything, and cuts 64/80 requests (512-token window; my prompts grew to 1–1.9k tokens with history)
    • Laya multilingual says "yes" to almost everything, including "normal" (0.97) to a night-time leak 💧
    • Jev in the house: VMC disagreements with the rule went from 100% to 6% once it got history. Half a cent per day.

    So the house stays on Jev. The local path is clear though: fine-tune on my own feedback labels, not on Jev's answers.

    Full post, commands, outputs and timings (EN/FR/IT):
    mornati.net/laya-vs-jev-ten-da

    Code & recipe: github.com/mmornati/system-one

    Thanks again to everyone who commented on the last Home Assistant post: the "why not a small local model?" question is what pushed me to measure this properly.

    #AI #LLM #FineTuning #HomeAssistant #SmartHome #AppleSilicon #OpenSource #Jev #Laya #SelfHosted

  7. New post: Laya vs Jev, ten days later 🧠⚖️

    Ten days ago I benchmarked two "System One" decision models for my LLM router: Jev (TypeSafe, hosted) was usable as-is, Laya (open weights) zero-shot was not. Laya's README says "fine-tune me", so I did, on a 16 GB M4 MacBook.

    🔁 Laya 0.3.6 → 0.3.24: 18 releases, but the weights are byte-identical. Same numbers.

    🆕 New open rivals on my 80-prompt bench:
    • Laya typed-decisions: 71% topic, but never confident → over-provisions everything
    • Von 1.3: closest to gold out of the box (48%), but over-confident
    • Kev-0.8B: 81% topic, 100% multilingual, very under-confident
    Jev 1.13: 89% topic, 72.5% same route as gold, still the only one usable as-is.

    🎯 Fine-tuning Laya on my own labels (not Jev's: distilling Jev is against TypeSafe's terms):
    • 480 labelled prompts → 1,920 training items
    • official Apple Silicon script, MPS
    • 41.5 min for 4 epochs, 15.4 GB peak memory
    • rule of thumb: ~6 s per labelled prompt

    Results:
    • topic 59% → 85%
    • complexity 45% → 72.5% (= Jev)
    • risk 31% → 69% (> Jev)
    • but 0% confident answers 🙃 → a temperature fitted on held-out data (T=0.87) fixes it
    • same route as gold: 20% → 64% (Jev 72.5%), routed cost $1.28 vs Jev's $1.35

    Caveats in the post: same labeller for train and test, tiny groups, and it under-provisions 9 prompts vs 3 for Jev.

    🏠 Then the reality check: I replayed 80 real Home Assistant decisions (VMC, shutters, water, watering) through Jev and stock Laya.
    • Laya English answers ~0.5 to everything, and cuts 64/80 requests (512-token window; my prompts grew to 1–1.9k tokens with history)
    • Laya multilingual says "yes" to almost everything, including "normal" (0.97) to a night-time leak 💧
    • Jev in the house: VMC disagreements with the rule went from 100% to 6% once it got history. Half a cent per day.

    So the house stays on Jev. The local path is clear though: fine-tune on my own feedback labels, not on Jev's answers.

    Full post, commands, outputs and timings (EN/FR/IT):
    mornati.net/laya-vs-jev-ten-da

    Code & recipe: github.com/mmornati/system-one

    Thanks again to everyone who commented on the last Home Assistant post: the "why not a small local model?" question is what pushed me to measure this properly.

    #AI #LLM #FineTuning #HomeAssistant #SmartHome #AppleSilicon #OpenSource #Jev #Laya #SelfHosted

  8. New post: Laya vs Jev, ten days later 🧠⚖️

    Ten days ago I benchmarked two "System One" decision models for my LLM router: Jev (TypeSafe, hosted) was usable as-is, Laya (open weights) zero-shot was not. Laya's README says "fine-tune me", so I did, on a 16 GB M4 MacBook.

    🔁 Laya 0.3.6 → 0.3.24: 18 releases, but the weights are byte-identical. Same numbers.

    🆕 New open rivals on my 80-prompt bench:
    • Laya typed-decisions: 71% topic, but never confident → over-provisions everything
    • Von 1.3: closest to gold out of the box (48%), but over-confident
    • Kev-0.8B: 81% topic, 100% multilingual, very under-confident
    Jev 1.13: 89% topic, 72.5% same route as gold, still the only one usable as-is.

    🎯 Fine-tuning Laya on my own labels (not Jev's: distilling Jev is against TypeSafe's terms):
    • 480 labelled prompts → 1,920 training items
    • official Apple Silicon script, MPS
    • 41.5 min for 4 epochs, 15.4 GB peak memory
    • rule of thumb: ~6 s per labelled prompt

    Results:
    • topic 59% → 85%
    • complexity 45% → 72.5% (= Jev)
    • risk 31% → 69% (> Jev)
    • but 0% confident answers 🙃 → a temperature fitted on held-out data (T=0.87) fixes it
    • same route as gold: 20% → 64% (Jev 72.5%), routed cost $1.28 vs Jev's $1.35

    Caveats in the post: same labeller for train and test, tiny groups, and it under-provisions 9 prompts vs 3 for Jev.

    🏠 Then the reality check: I replayed 80 real Home Assistant decisions (VMC, shutters, water, watering) through Jev and stock Laya.
    • Laya English answers ~0.5 to everything, and cuts 64/80 requests (512-token window; my prompts grew to 1–1.9k tokens with history)
    • Laya multilingual says "yes" to almost everything, including "normal" (0.97) to a night-time leak 💧
    • Jev in the house: VMC disagreements with the rule went from 100% to 6% once it got history. Half a cent per day.

    So the house stays on Jev. The local path is clear though: fine-tune on my own feedback labels, not on Jev's answers.

    Full post, commands, outputs and timings (EN/FR/IT):
    mornati.net/laya-vs-jev-ten-da

    Code & recipe: github.com/mmornati/system-one

    Thanks again to everyone who commented on the last Home Assistant post: the "why not a small local model?" question is what pushed me to measure this properly.

  9. RT @CloudflareDev: Wenn dir Jev gefällt, wirst du Clef und Clef-flash lieben – sie sind intelligenter, schneller und vollständig mit der Jev-API kompatibel.

    mehr auf Arint.info

    #API #Clef #Innovation #Jev #Software #Tech #arint_info

    https://x.com/CloudflareDev/status/2105780792559595666

  10. RT @CloudflareDev: Wenn dir Jev gefällt, wirst du Clef und Clef-flash lieben – sie sind intelligenter, schneller und vollständig mit der Jev-API kompatibel.

    mehr auf Arint.info

    #API #Clef #Innovation #Jev #Software #Tech #arint_info

    https://x.com/CloudflareDev/status/2105780792559595666

  11. Как Jev сэкономил нам 90% бюджета

    Как мы автоматизировали проверку новых товаров в интернет-магазине компьютерной техники, чтобы ликвидировать бэклог в 100 тыс. заявок. Отказ от стандартной LLM в пользу Jev.

    habr.com/ru/articles/1088934/

    #jev #ai #автоматизация #llm #crm

  12. Как Jev сэкономил нам 90% бюджета

    Как мы автоматизировали проверку новых товаров в интернет-магазине компьютерной техники, чтобы ликвидировать бэклог в 100 тыс. заявок. Отказ от стандартной LLM в пользу Jev.

    habr.com/ru/articles/1088934/

    #jev #ai #автоматизация #llm #crm

  13. Разбор Jev — модели, которая не умеет писать

    Утро началось с того, что ко мне в личку пришел друг и спросил: «Кто это такой, этот Jev? Ты работаешь с ИИ, должен знать». К этому моменту я уже видел десятки постов в ленте о том, насколько он крут, какие задачи решает, и призывы переносить все свои флоу на Jev. Давайте разбираться.

    habr.com/ru/companies/selectel

    #jev #ai #ml #jev_vs_llm #selectel #ии #ии_и_машинное_обучение

  14. Разбор Jev — модели, которая не умеет писать

    Утро началось с того, что ко мне в личку пришел друг и спросил: «Кто это такой, этот Jev? Ты работаешь с ИИ, должен знать». К этому моменту я уже видел десятки постов в ленте о том, насколько он крут, какие задачи решает, и призывы переносить все свои флоу на Jev. Давайте разбираться.

    habr.com/ru/companies/selectel

    #jev #ai #ml #jev_vs_llm #selectel #ии #ии_и_машинное_обучение

  15. It took plugins to make me realize the power of . 🤣 excellent overview on what you can do.

    Now scanning through a bunch of random local data to see of other quick analytical wins like this. 😁

    duckdb.org/2026/09/29/jev

  16. It took #duckdb plugins to make me realize the power of #jev. 🤣 excellent overview on what you can do.

    Now scanning through a bunch of random local data to see of other quick analytical wins like this. 😁

    duckdb.org/2026/09/29/jev

  17. It took #duckdb plugins to make me realize the power of #jev. 🤣 excellent overview on what you can do.

    Now scanning through a bunch of random local data to see of other quick analytical wins like this. 😁

    duckdb.org/2026/09/29/jev

  18. It took #duckdb plugins to make me realize the power of #jev. 🤣 excellent overview on what you can do.

    Now scanning through a bunch of random local data to see of other quick analytical wins like this. 😁

    duckdb.org/2026/09/29/jev

  19. Can Jev a super cheap, super fast classifier compete with LLMs for systematic review screening? aarontay.substack.com/p/can-jev-a-su… #AI #Jev

  20. Can Jev a super cheap, super fast classifier compete with LLMs for systematic review screening? aarontay.substack.com/p/can-jev-a-su… #AI #Jev

  21. Can Jev a super cheap, super fast classifier compete with LLMs for systematic review screening? aarontay.substack.com/p/can-jev-a-su… #AI #Jev

  22. Can Jev a super cheap, super fast classifier compete with LLMs for systematic review screening? aarontay.substack.com/p/can-jev-a-su… #AI #Jev

  23. Что умеет и где ломается Jev: большое тестирование

    В работе с моделями важно не только знать их точность, скорость и цену, но и понимать, где им можно доверять, а где нужна дополнительная проверка. Многие ограничения обнаруживаются уже в процессе интеграции — и о них хотелось бы знать заранее. Тестирование Jev я проводил прежде всего для себя и команды, но объём данных получился масштабным, и я решил поделиться результатами с сообществом. С 19 по 30 сентября 2026 года было проведено 22 эксперимента и отправлено около 160 тысяч запросов к jev-1.13.0 и контрольным моделям OpenAI. Все тесты запускались через API раннего доступа и суммарно обошлись примерно в $18 . Мне было интересно не столько проверить заявленные скорость и цену, сколько найти границы её практического применения. Насколько она точна по сравнению с небольшими GPT-моделями? Можно ли доверять её confidence и строить на нём автоматический роутинг? Что произойдёт при расширении списка классов, появлении длинного контекста, грязного входа или текста на другом языке? У LLM уже известны свои особенности — например, lost in the middle и склонность уверенно ошибаться. Логично было проверить, есть ли собственный набор таких эффектов у System One модели на примере Jev.

    habr.com/ru/articles/1089164/

    #jev #классификация_текста #языковые_модели #тестирование_моделей #калибровка_уверенности #zeroshot #обработка_естественного_языка #маршрутизация_запросов #llm

  24. Что умеет и где ломается Jev: большое тестирование

    В работе с моделями важно не только знать их точность, скорость и цену, но и понимать, где им можно доверять, а где нужна дополнительная проверка. Многие ограничения обнаруживаются уже в процессе интеграции — и о них хотелось бы знать заранее. Тестирование Jev я проводил прежде всего для себя и команды, но объём данных получился масштабным, и я решил поделиться результатами с сообществом. С 19 по 30 сентября 2026 года было проведено 22 эксперимента и отправлено около 160 тысяч запросов к jev-1.13.0 и контрольным моделям OpenAI. Все тесты запускались через API раннего доступа и суммарно обошлись примерно в $18 . Мне было интересно не столько проверить заявленные скорость и цену, сколько найти границы её практического применения. Насколько она точна по сравнению с небольшими GPT-моделями? Можно ли доверять её confidence и строить на нём автоматический роутинг? Что произойдёт при расширении списка классов, появлении длинного контекста, грязного входа или текста на другом языке? У LLM уже известны свои особенности — например, lost in the middle и склонность уверенно ошибаться. Логично было проверить, есть ли собственный набор таких эффектов у System One модели на примере Jev.

    habr.com/ru/articles/1089164/

    #jev #классификация_текста #языковые_модели #тестирование_моделей #калибровка_уверенности #zeroshot #обработка_естественного_языка #маршрутизация_запросов #llm