Topic Hub

AI & ML on Cloud Native

Running AI/ML workloads on Kubernetes and modern infrastructure.

GPU scheduling, model serving, Kubeflow, LLMs on Kubernetes, NVIDIA NVCF, and the rapidly evolving AI infrastructure landscape.

Start here

More on AI & ML on Cloud Native

Cover for Running a big LLM across multiple GPUs with vLLM
vllmgpuAug 18, 2026

Running a big LLM across multiple GPUs with vLLM

A plain-English guide to serving a model too big for one GPU, in two tracks: a runbook from download to serving with every flag and error explained, and a deep dive into how tensor, pipeline, and expert parallelism split the model, with measured numbers from a 235B model on four RTX PRO 6000 cards.

Shubham KataraSaiyam Pathak
Shubham Katara & Saiyam Pathak · 53 min
Read →
Cover for The Local LLM Glossary: Every Term, Flag, and Number in Plain English
local-aillmAug 18, 2026

The Local LLM Glossary: Every Term, Flag, and Number in Plain English

Plain-English definitions for every term you hit in local LLM posts: prefill and decode, tokens per second, FP8 and NVFP4, Q4_K_M, KV cache, YaRN, Gated DeltaNet, speculative decoding, and every vLLM, llama.cpp, and Ollama flag worth knowing.

Saiyam Pathak
Saiyam Pathak · 19 min
Read →
Cover for Running Qwen3.8-27B on DGX Spark
qwendgxsparkAug 17, 2026

Running Qwen3.8-27B on DGX Spark

Qwen3.8-27B on DGX Spark with llama.cpp, Ollama, vLLM, and SGLang: the recipes, the tokens per second I measured, MTP speculative decoding, and the sharp edges I hit along the way.

Saiyam Pathak
Saiyam Pathak · 24 min
Read →
Cover for I Ran an AI SRE Copilot on My Own Hardware. Here Is What It Actually Does.
kubernetesaiAug 17, 2026

I Ran an AI SRE Copilot on My Own Hardware. Here Is What It Actually Does.

Running NudgeBee v1.4.0 end to end - a self-hosted AIOps platform behind AI-SRE, AI-FinOps, AI-K8sOps, and agentic automation - on a Mac, a kiac cluster, and a DGX Spark.

Saiyam Pathak
Saiyam Pathak · 21 min
Read →
Cover for Running Nemotron 3.5 Lightning on DGX Spark
nvidiadgxsparkAug 11, 2026

Running Nemotron 3.5 Lightning on DGX Spark

NVIDIA's new Nemotron 3.5 Lightning on DGX Spark: how to run it with Ollama and vLLM, the tokens per second I measured, and how the two paths compare.

Saiyam Pathak
Saiyam Pathak · 13 min
Read →
Cover for HAMi Dynamic MIG on RTX PRO 6000: A Live Kubernetes Test
kubernetesgpuAug 11, 2026

HAMi Dynamic MIG on RTX PRO 6000: A Live Kubernetes Test

Hands-on HAMi Dynamic MIG test on Kubernetes and RTX PRO 6000 Blackwell: setup commands, real allocations, mixed profiles, reclamation, and recovery.

Shubham KataraSaiyam Pathak
Shubham Katara & Saiyam Pathak · 24 min
Read →
Cover for How to Share GPUs in Kubernetes at Scale with HAMi (Software vGPU Slicing)
kubernetesgpuJul 23, 2026

How to Share GPUs in Kubernetes at Scale with HAMi (Software vGPU Slicing)

Share NVIDIA GPUs in Kubernetes with HAMi software vGPU slicing: memory and compute limits, Helm configuration, a verified PyTorch manifest, a real RTX PRO 6000 OOM test, and Prometheus monitoring.

Shubham KataraSaiyam Pathak
Shubham Katara & Saiyam Pathak · 34 min
Read →
Cover for Slicing GPUs in Kubernetes with NVIDIA Multi-Instance GPU (MIG)
kubernetesgpuJul 20, 2026

Slicing GPUs in Kubernetes with NVIDIA Multi-Instance GPU (MIG)

GPU sharing in Kubernetes explained: time-slicing vs MPS vs MIG, every nvidia-smi command to enable and disable MIG on one GPU or eight, GPU Operator automation, pitfalls, and DCGM monitoring.

Shubham KataraSaiyam Pathak
Shubham Katara & Saiyam Pathak · 45 min
Read →
Cover for Day 5: Local LLM Inference Engines, Wrappers, and What to Pick
nvidiadgxsparkJul 17, 2026

Day 5: Local LLM Inference Engines, Wrappers, and What to Pick

A beginner-friendly guide to local LLM inference, with the same Qwen model tested through Ollama, llama.cpp, Docker Model Runner, vLLM, SGLang, and TensorRT-LLM on NVIDIA DGX Spark.

Saiyam Pathak
Saiyam Pathak · 56 min
Read →
Cover for Bonsai 27B on RTX PRO 6000 vs DGX Spark: what actually works
aigpuJul 16, 2026

Bonsai 27B on RTX PRO 6000 vs DGX Spark: what actually works

Real Bonsai 27B benchmarks on an RTX PRO 6000 and a DGX Spark, including the supported llama.cpp setup, ternary vs 1-bit results, and speculative decoding.

Saiyam Pathak
Saiyam Pathak · 14 min
Read →
Cover for Day 4: Quantization Demystified. BF16, FP8, NVFP4, MXFP4, INT4, GGUF, and Why It All Matters
nvidiadgxsparkJun 10, 2026

Day 4: Quantization Demystified. BF16, FP8, NVFP4, MXFP4, INT4, GGUF, and Why It All Matters

A practical, beginner-friendly guide to BF16, FP8, NVFP4, MXFP4, INT4, and GGUF Q4_K_M on NVIDIA DGX Spark. Bytes per parameter, quality vs size, and which format to pick when.

Saiyam Pathak
Saiyam Pathak · 28 min
Read →
Cover for Day 3: The DGX Spark Unpacked. GB10, Unified Memory, sm_121, and the One Reason This Hardware Exists
nvidiadgxsparkJun 5, 2026

Day 3: The DGX Spark Unpacked. GB10, Unified Memory, sm_121, and the One Reason This Hardware Exists

A practical teardown of NVIDIA DGX Spark's GB10 Grace Blackwell Superchip, unified memory, sm_121, NVFP4 tensor cores, memory reporting, and decode limits.

Saiyam Pathak
Saiyam Pathak · 19 min
Read →
Show 11 more AI & ML on Cloud Native articles