home.social

#llamastack — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #llamastack, aggregated by home.social.

fetched live
  1. As LLMs move into production, #Observability is essential for Reliability, Performance & Responsible AI.

    Learn how to deploy an #opensource observability stack - using Prometheus, Grafana, Tempo, and OpenTelemetry Collectors on Kubernetes - and monitor real #AI workloads with #vLLM & #Llamastack.

    🎥 Watch the #InfoQ video (#transcript included): bit.ly/4hlKoDa

    #Prometheus #Grafana #OpenTelemetry #Kubernetes

  2. As LLMs move into production, is essential for Reliability, Performance & Responsible AI.

    Learn how to deploy an observability stack - using Prometheus, Grafana, Tempo, and OpenTelemetry Collectors on Kubernetes - and monitor real workloads with & .

    🎥 Watch the video (#transcript included): bit.ly/4hlKoDa

  3. 🦙 #LlamaStack: Standardizing #GenerativeAI Development

    Defines open API specs for #AI application building blocks

    Covers full lifecycle: model training, evaluation, production deployment

    Includes APIs for inference, safety, memory, agents, and more

    Supports multiple environments: local, hosted, and on-device

    🛠️ Features:

    #OpenSource API providers and distributions

    Mix-and-match capabilities (e.g., local small models, cloud-based large models)

    Consistent APIs across platforms (server, mobile, etc.)

    🤝 Supported implementations:

    API Providers: #Meta Reference, #Fireworks, #AWS Bedrock, #Together, #Ollama, TGI, #Chroma, PG Vector, #PyTorch ExecuTorch

    Distributions: Meta Reference, Dell-TGI

    📦 Easy installation via pip or from source 🖥️ Includes 'llama' CLI for managing distributions, models, and more

    Learn more: github.com/meta-llama/llama-st

  4. 🦙 #LlamaStack: Standardizing #GenerativeAI Development

    Defines open API specs for #AI application building blocks

    Covers full lifecycle: model training, evaluation, production deployment

    Includes APIs for inference, safety, memory, agents, and more

    Supports multiple environments: local, hosted, and on-device

    🛠️ Features:

    #OpenSource API providers and distributions

    Mix-and-match capabilities (e.g., local small models, cloud-based large models)

    Consistent APIs across platforms (server, mobile, etc.)

    🤝 Supported implementations:

    API Providers: #Meta Reference, #Fireworks, #AWS Bedrock, #Together, #Ollama, TGI, #Chroma, PG Vector, #PyTorch ExecuTorch

    Distributions: Meta Reference, Dell-TGI

    📦 Easy installation via pip or from source 🖥️ Includes 'llama' CLI for managing distributions, models, and more

    Learn more: github.com/meta-llama/llama-st