Close Menu
    Facebook X (Twitter) Instagram
    • Privacy Policy
    • Terms Of Service
    • Social Media Disclaimer
    • DMCA Compliance
    • Anti-Spam Policy
    Facebook X (Twitter) Instagram
    Bytecore News
    • Home
    • Crypto News
      • Bitcoin
      • Ethereum
      • Altcoins
      • Blockchain
      • DeFi
    • AI News
    • Stock News
    • Learn
      • AI for Beginners
      • AI Tips
      • Make Money with AI
    • Reviews
    • Tools
      • Best AI Tools
      • Crypto Market Cap List
      • Stock Market Overview
      • Market Heatmap
    • Contact
    Bytecore News
    Home»AI News»Microsoft Open-Sources TauGrid: A Kubernetes-Native Stack for GPU AI Workloads
    Microsoft Open-Sources TauGrid: A Kubernetes-Native Stack for GPU AI Workloads
    AI News

    Microsoft Open-Sources TauGrid: A Kubernetes-Native Stack for GPU AI Workloads

    September 18, 20264 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email
    Customgpt


    Platform teams running AI on Kubernetes rarely run one thing. They run a queueing system, a distributed runtime, GPU node health checks, dashboards, and a layer of submission scripts holding all of it together. The Azure Kubernetes Service engineering team open-sourced TauGrid, which collapses that assembly job into a single Helm install.

    Is it deployable? Yes, TauGrid is MIT licensed, with container images and Helm charts published as public OCI artifacts on Microsoft Container Registry. Prerequisites are a Kubernetes 1.30+ cluster with GPU nodes, kubectl, and Helm 3.0 or later.

    What is TauGrid

    TauGrid is a self-hosted platform for running AI workloads on Kubernetes. It combines five things that platform teams usually integrate by hand: the tau CLI, workload queueing and admission through Kueue, Ray cluster orchestration through KubeRay, node-level GPU health monitoring, and cluster and workload observability.

    The split of responsibility is the design point. Platform teams own workspaces, queues, compute profiles, storage, identity, and observability. Researchers work from a repository and the CLI, and submit workloads without configuring Kubernetes directly. The codebase is written primarily in Go.

    kraken

    How a job moves through it

    A workload is described in a tau.yaml file. The GPU training example published by Microsoft runs a PyTorch job on a single A100:

    schema_version: 1
    name: aks-gpu-quickstart
    run:
    entrypoint: train.py
    workload_kind: rayjob
    compute:
    gpus: 1
    workers: 1
    cpus: 16
    memory: 64Gi
    runtime:
    image: mcr.microsoft.com/aks/ai-runtime/ray:py3.12-ray2.56.0-cuda13.0
    pip:
    – torch>=2.4.0

    On tau run, TauGrid resolves platform policy, renders a Kubernetes Job or a KubeRay RayJob, and submits it through Kueue. The six stages Microsoft documents are submission, queueing, execution, monitoring, recovery, and evidence. Recovery covers retry, resume from checkpoint, and failure diagnosis. Evidence records capture workload metadata, configuration, logs, metrics, checkpoints, and execution history, which is what makes a run reproducible and auditable later.

    When several teams share a cluster, their jobs land in a shared Kueue ClusterQueue. Kueue admits each one on quota and priority, and Kubernetes places it on healthy GPUs.

    Interactive explainer

    Installation is a Helm chart pulled straight from MCR:

    helm install taugrid \
    oci://mcr.microsoft.com/aks/ai-runtime/helm/taugrid \
    –version 0.4.2 \
    –namespace tau-system \
    –create-namespace

    First-party images ship under mcr.microsoft.com/aks/ai-runtime/ for Tau, the TauGrid Portal, and the tau core controller. Microsoft advises pinning versioned tags or immutable digests rather than latest. The CLI installs from GitHub Releases on Linux and macOS, with a PowerShell installer for Windows amd64; the installer verifies the release checksum and does not modify PATH.

    Two operational details matter for anyone evaluating this outside Azure. First, TauGrid sends no telemetry to Microsoft by default, and remote export stays off unless an operator configures a destination. Second, some integrations are still Azure-specific, notably observability through Azure Data Explorer. The stated intent is to support cloud and on-premises Kubernetes without an Azure dependency, and contributions toward that are open.

    Key Takeaways

    • Microsoft open-sourced TauGrid on August 28, 2026, under the MIT license at Azure/taugrid.
    • One Helm install bundles the tau CLI, Kueue queueing, KubeRay orchestration, GPU health monitoring, and observability.
    • Deployable now on any Kubernetes 1.30+ cluster with GPU nodes, kubectl, and Helm 3.0+.
    • Evidence records capture config, logs, metrics, and checkpoints, so runs stay reproducible and auditable.
    • No telemetry by default, but Azure Data Explorer observability remains Azure-specific for now.

    Check out the AKS Engineering Blog and Azure/taugrid on GitHub. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

    Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us

    Asif Razzaq is the CEO of Marktechpost AI Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.



    Source link

    frase
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    CryptoExpert
    • Website

    Related Posts

    New AI technique could make minimally invasive surgeries safer and more precise | MIT News

    September 17, 2026

    Pony.ai unveils autonomous electric truck for logistics fleets

    September 16, 2026

    Agent Harness vs Agent Framework vs MCP: Which Layer Owns the Loop, State, Tools, Permissions, and Recovery

    September 15, 2026

    Lifesaving Lincoln Laboratory device wins 2026 Excellence in Technology Transfer Award | MIT News

    September 14, 2026
    Add A Comment
    Leave A Reply Cancel Reply

    binance
    Latest Posts

    Cattle Fall Back as Another Border Opening Announced

    September 18, 2026

    Microsoft Open-Sources TauGrid: A Kubernetes-Native Stack for GPU AI Workloads

    September 18, 2026

    Making Money With AI Is On “Easy Mode” For Beginners in 2026

    September 18, 2026

    Claude AI Complete Course for Beginners 2026 – Start Here

    September 17, 2026

    Generative AI 101: Explained in 14 Minutes!

    September 17, 2026
    coinbase
    LEGAL INFORMATION
    • Privacy Policy
    • Terms Of Service
    • Social Media Disclaimer
    • DMCA Compliance
    • Anti-Spam Policy
    Top Insights

    Bitcoin Gains New Bull Signal as Fisher Transform Prints Key Crossover

    September 18, 2026

    Dragonfly’s Qureshi Calls for End to Zcash Dev Fund After 2028

    September 18, 2026
    Customgpt
    Facebook X (Twitter) Instagram Pinterest
    © 2026 BytecoreNews.com - All rights reserved.

    Type above and press Enter to search. Press Esc to cancel.