#jev — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #jev, aggregated by home.social.
-
#Cloudflare released #Clef and #Clefflash #decisionmodels, claiming faster, more accurate performance than #Jev via #Qwen backbones. Clef handles images and video, supports 64k context, costs $0.24 per million tokens, and is Jev-compatible on Workers AI or Hugging Face. https://www.theregister.com/ai-and-ml/2026/10/01/cloudflare-tries-to-outplay-jev-with-open-weight-clef-models/5300649?eicker.news #tech #news #ainews
-
#Cloudflare released #Clef and #Clefflash #decisionmodels, claiming faster, more accurate performance than #Jev via #Qwen backbones. Clef handles images and video, supports 64k context, costs $0.24 per million tokens, and is Jev-compatible on Workers AI or Hugging Face. https://www.theregister.com/ai-and-ml/2026/10/01/cloudflare-tries-to-outplay-jev-with-open-weight-clef-models/5300649?eicker.news #tech #news #ainews
-
#Cloudflare released #Clef and #Clefflash #decisionmodels, claiming faster, more accurate performance than #Jev via #Qwen backbones. Clef handles images and video, supports 64k context, costs $0.24 per million tokens, and is Jev-compatible on Workers AI or Hugging Face. https://www.theregister.com/ai-and-ml/2026/10/01/cloudflare-tries-to-outplay-jev-with-open-weight-clef-models/5300649?eicker.news #tech #news #ainews
-
#Cloudflare released #Clef and #Clefflash #decisionmodels, claiming faster, more accurate performance than #Jev via #Qwen backbones. Clef handles images and video, supports 64k context, costs $0.24 per million tokens, and is Jev-compatible on Workers AI or Hugging Face. https://www.theregister.com/ai-and-ml/2026/10/01/cloudflare-tries-to-outplay-jev-with-open-weight-clef-models/5300649?eicker.news #tech #news #ainews
-
New post: Laya vs Jev, ten days later 🧠⚖️
Ten days ago I benchmarked two "System One" decision models for my LLM router: Jev (TypeSafe, hosted) was usable as-is, Laya (open weights) zero-shot was not. Laya's README says "fine-tune me", so I did, on a 16 GB M4 MacBook.
🔁 Laya 0.3.6 → 0.3.24: 18 releases, but the weights are byte-identical. Same numbers.
🆕 New open rivals on my 80-prompt bench:
• Laya typed-decisions: 71% topic, but never confident → over-provisions everything
• Von 1.3: closest to gold out of the box (48%), but over-confident
• Kev-0.8B: 81% topic, 100% multilingual, very under-confident
Jev 1.13: 89% topic, 72.5% same route as gold, still the only one usable as-is.🎯 Fine-tuning Laya on my own labels (not Jev's: distilling Jev is against TypeSafe's terms):
• 480 labelled prompts → 1,920 training items
• official Apple Silicon script, MPS
• 41.5 min for 4 epochs, 15.4 GB peak memory
• rule of thumb: ~6 s per labelled promptResults:
• topic 59% → 85%
• complexity 45% → 72.5% (= Jev)
• risk 31% → 69% (> Jev)
• but 0% confident answers 🙃 → a temperature fitted on held-out data (T=0.87) fixes it
• same route as gold: 20% → 64% (Jev 72.5%), routed cost $1.28 vs Jev's $1.35Caveats in the post: same labeller for train and test, tiny groups, and it under-provisions 9 prompts vs 3 for Jev.
🏠 Then the reality check: I replayed 80 real Home Assistant decisions (VMC, shutters, water, watering) through Jev and stock Laya.
• Laya English answers ~0.5 to everything, and cuts 64/80 requests (512-token window; my prompts grew to 1–1.9k tokens with history)
• Laya multilingual says "yes" to almost everything, including "normal" (0.97) to a night-time leak 💧
• Jev in the house: VMC disagreements with the rule went from 100% to 6% once it got history. Half a cent per day.So the house stays on Jev. The local path is clear though: fine-tune on my own feedback labels, not on Jev's answers.
Full post, commands, outputs and timings (EN/FR/IT):
https://mornati.net/laya-vs-jev-ten-days-later-fine-tuning-on-a-macbook/Code & recipe: https://github.com/mmornati/system-one-router
Thanks again to everyone who commented on the last Home Assistant post: the "why not a small local model?" question is what pushed me to measure this properly.
#AI #LLM #FineTuning #HomeAssistant #SmartHome #AppleSilicon #OpenSource #Jev #Laya #SelfHosted
-
New post: Laya vs Jev, ten days later 🧠⚖️
Ten days ago I benchmarked two "System One" decision models for my LLM router: Jev (TypeSafe, hosted) was usable as-is, Laya (open weights) zero-shot was not. Laya's README says "fine-tune me", so I did, on a 16 GB M4 MacBook.
🔁 Laya 0.3.6 → 0.3.24: 18 releases, but the weights are byte-identical. Same numbers.
🆕 New open rivals on my 80-prompt bench:
• Laya typed-decisions: 71% topic, but never confident → over-provisions everything
• Von 1.3: closest to gold out of the box (48%), but over-confident
• Kev-0.8B: 81% topic, 100% multilingual, very under-confident
Jev 1.13: 89% topic, 72.5% same route as gold, still the only one usable as-is.🎯 Fine-tuning Laya on my own labels (not Jev's: distilling Jev is against TypeSafe's terms):
• 480 labelled prompts → 1,920 training items
• official Apple Silicon script, MPS
• 41.5 min for 4 epochs, 15.4 GB peak memory
• rule of thumb: ~6 s per labelled promptResults:
• topic 59% → 85%
• complexity 45% → 72.5% (= Jev)
• risk 31% → 69% (> Jev)
• but 0% confident answers 🙃 → a temperature fitted on held-out data (T=0.87) fixes it
• same route as gold: 20% → 64% (Jev 72.5%), routed cost $1.28 vs Jev's $1.35Caveats in the post: same labeller for train and test, tiny groups, and it under-provisions 9 prompts vs 3 for Jev.
🏠 Then the reality check: I replayed 80 real Home Assistant decisions (VMC, shutters, water, watering) through Jev and stock Laya.
• Laya English answers ~0.5 to everything, and cuts 64/80 requests (512-token window; my prompts grew to 1–1.9k tokens with history)
• Laya multilingual says "yes" to almost everything, including "normal" (0.97) to a night-time leak 💧
• Jev in the house: VMC disagreements with the rule went from 100% to 6% once it got history. Half a cent per day.So the house stays on Jev. The local path is clear though: fine-tune on my own feedback labels, not on Jev's answers.
Full post, commands, outputs and timings (EN/FR/IT):
https://mornati.net/laya-vs-jev-ten-days-later-fine-tuning-on-a-macbook/Code & recipe: https://github.com/mmornati/system-one-router
Thanks again to everyone who commented on the last Home Assistant post: the "why not a small local model?" question is what pushed me to measure this properly.
#AI #LLM #FineTuning #HomeAssistant #SmartHome #AppleSilicon #OpenSource #Jev #Laya #SelfHosted
-
New post: Laya vs Jev, ten days later 🧠⚖️
Ten days ago I benchmarked two "System One" decision models for my LLM router: Jev (TypeSafe, hosted) was usable as-is, Laya (open weights) zero-shot was not. Laya's README says "fine-tune me", so I did, on a 16 GB M4 MacBook.
🔁 Laya 0.3.6 → 0.3.24: 18 releases, but the weights are byte-identical. Same numbers.
🆕 New open rivals on my 80-prompt bench:
• Laya typed-decisions: 71% topic, but never confident → over-provisions everything
• Von 1.3: closest to gold out of the box (48%), but over-confident
• Kev-0.8B: 81% topic, 100% multilingual, very under-confident
Jev 1.13: 89% topic, 72.5% same route as gold, still the only one usable as-is.🎯 Fine-tuning Laya on my own labels (not Jev's: distilling Jev is against TypeSafe's terms):
• 480 labelled prompts → 1,920 training items
• official Apple Silicon script, MPS
• 41.5 min for 4 epochs, 15.4 GB peak memory
• rule of thumb: ~6 s per labelled promptResults:
• topic 59% → 85%
• complexity 45% → 72.5% (= Jev)
• risk 31% → 69% (> Jev)
• but 0% confident answers 🙃 → a temperature fitted on held-out data (T=0.87) fixes it
• same route as gold: 20% → 64% (Jev 72.5%), routed cost $1.28 vs Jev's $1.35Caveats in the post: same labeller for train and test, tiny groups, and it under-provisions 9 prompts vs 3 for Jev.
🏠 Then the reality check: I replayed 80 real Home Assistant decisions (VMC, shutters, water, watering) through Jev and stock Laya.
• Laya English answers ~0.5 to everything, and cuts 64/80 requests (512-token window; my prompts grew to 1–1.9k tokens with history)
• Laya multilingual says "yes" to almost everything, including "normal" (0.97) to a night-time leak 💧
• Jev in the house: VMC disagreements with the rule went from 100% to 6% once it got history. Half a cent per day.So the house stays on Jev. The local path is clear though: fine-tune on my own feedback labels, not on Jev's answers.
Full post, commands, outputs and timings (EN/FR/IT):
https://mornati.net/laya-vs-jev-ten-days-later-fine-tuning-on-a-macbook/Code & recipe: https://github.com/mmornati/system-one-router
Thanks again to everyone who commented on the last Home Assistant post: the "why not a small local model?" question is what pushed me to measure this properly.
#AI #LLM #FineTuning #HomeAssistant #SmartHome #AppleSilicon #OpenSource #Jev #Laya #SelfHosted
-
New post: Laya vs Jev, ten days later 🧠⚖️
Ten days ago I benchmarked two "System One" decision models for my LLM router: Jev (TypeSafe, hosted) was usable as-is, Laya (open weights) zero-shot was not. Laya's README says "fine-tune me", so I did, on a 16 GB M4 MacBook.
🔁 Laya 0.3.6 → 0.3.24: 18 releases, but the weights are byte-identical. Same numbers.
🆕 New open rivals on my 80-prompt bench:
• Laya typed-decisions: 71% topic, but never confident → over-provisions everything
• Von 1.3: closest to gold out of the box (48%), but over-confident
• Kev-0.8B: 81% topic, 100% multilingual, very under-confident
Jev 1.13: 89% topic, 72.5% same route as gold, still the only one usable as-is.🎯 Fine-tuning Laya on my own labels (not Jev's: distilling Jev is against TypeSafe's terms):
• 480 labelled prompts → 1,920 training items
• official Apple Silicon script, MPS
• 41.5 min for 4 epochs, 15.4 GB peak memory
• rule of thumb: ~6 s per labelled promptResults:
• topic 59% → 85%
• complexity 45% → 72.5% (= Jev)
• risk 31% → 69% (> Jev)
• but 0% confident answers 🙃 → a temperature fitted on held-out data (T=0.87) fixes it
• same route as gold: 20% → 64% (Jev 72.5%), routed cost $1.28 vs Jev's $1.35Caveats in the post: same labeller for train and test, tiny groups, and it under-provisions 9 prompts vs 3 for Jev.
🏠 Then the reality check: I replayed 80 real Home Assistant decisions (VMC, shutters, water, watering) through Jev and stock Laya.
• Laya English answers ~0.5 to everything, and cuts 64/80 requests (512-token window; my prompts grew to 1–1.9k tokens with history)
• Laya multilingual says "yes" to almost everything, including "normal" (0.97) to a night-time leak 💧
• Jev in the house: VMC disagreements with the rule went from 100% to 6% once it got history. Half a cent per day.So the house stays on Jev. The local path is clear though: fine-tune on my own feedback labels, not on Jev's answers.
Full post, commands, outputs and timings (EN/FR/IT):
https://mornati.net/laya-vs-jev-ten-days-later-fine-tuning-on-a-macbook/Code & recipe: https://github.com/mmornati/system-one-router
Thanks again to everyone who commented on the last Home Assistant post: the "why not a small local model?" question is what pushed me to measure this properly.
#AI #LLM #FineTuning #HomeAssistant #SmartHome #AppleSilicon #OpenSource #Jev #Laya #SelfHosted
-
Claude Code ModsとJevでモデルルーターを構築(サブスクで使えるゾ!!)
https://qiita.com/moritalous/items/8b663db633dde3c49d62?utm_campaign=popular_items&utm_medium=feed&utm_source=popular_items -
#llamacpp now supports local #jev like decision models https://huggingface.co/blog/ggml-org/decision-models-in-llamacpp #ai #llm
-
#llamacpp now supports local #jev like decision models https://huggingface.co/blog/ggml-org/decision-models-in-llamacpp #ai #llm
-
#llamacpp now supports local #jev like decision models https://huggingface.co/blog/ggml-org/decision-models-in-llamacpp #ai #llm
-
#llamacpp now supports local #jev like decision models https://huggingface.co/blog/ggml-org/decision-models-in-llamacpp #ai #llm
-
#Design #Explainers
Jev for designers · Why the AI model is different and why it matters https://ilo.im/16gmud_____
#Jev #AiModels #AI #Decisions #ProductDesign #UxDesign #UiDesign #WebDesign -
#Design #Explainers
Jev for designers · Why the AI model is different and why it matters https://ilo.im/16gmud_____
#Jev #AiModels #AI #Decisions #ProductDesign #UxDesign #UiDesign #WebDesign -
#Design #Explainers
Jev for designers · Why the AI model is different and why it matters https://ilo.im/16gmud_____
#Jev #AiModels #AI #Decisions #ProductDesign #UxDesign #UiDesign #WebDesign -
RT @CloudflareDev: Wenn dir Jev gefällt, wirst du Clef und Clef-flash lieben – sie sind intelligenter, schneller und vollständig mit der Jev-API kompatibel.
mehr auf Arint.info
-
RT @CloudflareDev: Wenn dir Jev gefällt, wirst du Clef und Clef-flash lieben – sie sind intelligenter, schneller und vollständig mit der Jev-API kompatibel.
mehr auf Arint.info
-
Как Jev сэкономил нам 90% бюджета
Как мы автоматизировали проверку новых товаров в интернет-магазине компьютерной техники, чтобы ликвидировать бэклог в 100 тыс. заявок. Отказ от стандартной LLM в пользу Jev.
-
Как Jev сэкономил нам 90% бюджета
Как мы автоматизировали проверку новых товаров в интернет-магазине компьютерной техники, чтобы ликвидировать бэклог в 100 тыс. заявок. Отказ от стандартной LLM в пользу Jev.
-
Разбор Jev — модели, которая не умеет писать
Утро началось с того, что ко мне в личку пришел друг и спросил: «Кто это такой, этот Jev? Ты работаешь с ИИ, должен знать». К этому моменту я уже видел десятки постов в ленте о том, насколько он крут, какие задачи решает, и призывы переносить все свои флоу на Jev. Давайте разбираться.
https://habr.com/ru/companies/selectel/articles/1089416/
#jev #ai #ml #jev_vs_llm #selectel #ии #ии_и_машинное_обучение
-
Разбор Jev — модели, которая не умеет писать
Утро началось с того, что ко мне в личку пришел друг и спросил: «Кто это такой, этот Jev? Ты работаешь с ИИ, должен знать». К этому моменту я уже видел десятки постов в ленте о том, насколько он крут, какие задачи решает, и призывы переносить все свои флоу на Jev. Давайте разбираться.
https://habr.com/ru/companies/selectel/articles/1089416/
#jev #ai #ml #jev_vs_llm #selectel #ии #ии_и_машинное_обучение
-
Can Jev a super cheap, super fast classifier compete with LLMs for systematic review screening? aarontay.substack.com/p/can-jev-a-su… #AI #Jev
-
Can Jev a super cheap, super fast classifier compete with LLMs for systematic review screening? aarontay.substack.com/p/can-jev-a-su… #AI #Jev
-
Can Jev a super cheap, super fast classifier compete with LLMs for systematic review screening? aarontay.substack.com/p/can-jev-a-su… #AI #Jev
-
Can Jev a super cheap, super fast classifier compete with LLMs for systematic review screening? aarontay.substack.com/p/can-jev-a-su… #AI #Jev
-
-
-
-
-
Что умеет и где ломается Jev: большое тестирование
В работе с моделями важно не только знать их точность, скорость и цену, но и понимать, где им можно доверять, а где нужна дополнительная проверка. Многие ограничения обнаруживаются уже в процессе интеграции — и о них хотелось бы знать заранее. Тестирование Jev я проводил прежде всего для себя и команды, но объём данных получился масштабным, и я решил поделиться результатами с сообществом. С 19 по 30 сентября 2026 года было проведено 22 эксперимента и отправлено около 160 тысяч запросов к jev-1.13.0 и контрольным моделям OpenAI. Все тесты запускались через API раннего доступа и суммарно обошлись примерно в $18 . Мне было интересно не столько проверить заявленные скорость и цену, сколько найти границы её практического применения. Насколько она точна по сравнению с небольшими GPT-моделями? Можно ли доверять её confidence и строить на нём автоматический роутинг? Что произойдёт при расширении списка классов, появлении длинного контекста, грязного входа или текста на другом языке? У LLM уже известны свои особенности — например, lost in the middle и склонность уверенно ошибаться. Логично было проверить, есть ли собственный набор таких эффектов у System One модели на примере Jev.
https://habr.com/ru/articles/1089164/
#jev #классификация_текста #языковые_модели #тестирование_моделей #калибровка_уверенности #zeroshot #обработка_естественного_языка #маршрутизация_запросов #llm
-
Что умеет и где ломается Jev: большое тестирование
В работе с моделями важно не только знать их точность, скорость и цену, но и понимать, где им можно доверять, а где нужна дополнительная проверка. Многие ограничения обнаруживаются уже в процессе интеграции — и о них хотелось бы знать заранее. Тестирование Jev я проводил прежде всего для себя и команды, но объём данных получился масштабным, и я решил поделиться результатами с сообществом. С 19 по 30 сентября 2026 года было проведено 22 эксперимента и отправлено около 160 тысяч запросов к jev-1.13.0 и контрольным моделям OpenAI. Все тесты запускались через API раннего доступа и суммарно обошлись примерно в $18 . Мне было интересно не столько проверить заявленные скорость и цену, сколько найти границы её практического применения. Насколько она точна по сравнению с небольшими GPT-моделями? Можно ли доверять её confidence и строить на нём автоматический роутинг? Что произойдёт при расширении списка классов, появлении длинного контекста, грязного входа или текста на другом языке? У LLM уже известны свои особенности — например, lost in the middle и склонность уверенно ошибаться. Логично было проверить, есть ли собственный набор таких эффектов у System One модели на примере Jev.
https://habr.com/ru/articles/1089164/
#jev #классификация_текста #языковые_модели #тестирование_моделей #калибровка_уверенности #zeroshot #обработка_естественного_языка #маршрутизация_запросов #llm
-
Jev のような判定モデルを SQL から呼べる ai_decide を日本語で試した
https://qiita.com/taka_yayoi/items/5e5b4baa7781ff6008bd?utm_campaign=popular_items&utm_medium=feed&utm_source=popular_items