Nouveauté

Chaos Engineering for AI Infrastructure

Par : Neel Trove
Offrir maintenant
Ou planifier dans votre panier
Disponible dans votre compte client Decitre ou Furet du Nord dès validation de votre commande. Le format ePub protégé est :
  • Compatible avec une lecture sur My Vivlio (smartphone, tablette, ordinateur)
  • Compatible avec une lecture sur liseuses Vivlio
  • Pour les liseuses autres que Vivlio, vous devez utiliser le logiciel Adobe Digital Edition. Non compatible avec la lecture sur les liseuses Kindle, Remarkable et Sony
  • Non compatible avec un achat hors France métropolitaine
Logo Vivlio, qui est-ce ?

Notre partenaire de plateforme de lecture numérique où vous retrouverez l'ensemble de vos ebooks gratuitement

Pour en savoir plus sur nos ebooks, consultez notre aide en ligne ici
C'est si simple ! Lisez votre ebook avec l'app Vivlio sur votre tablette, mobile ou ordinateur :
Google PlayApp Store
  • FormatePub
  • ISBN8235980051
  • EAN9798235980051
  • Date de parution19/08/2026
  • Protection num.Adobe DRM
  • Infos supplémentairesepub
  • ÉditeurIoakim Ioakim

Résumé

How long do you think it would take you to notice if your AI system started giving worse answers?Honestly, for most teams, it's weeks, and the news comes from a customer. Chaos engineering answers that question for infrastructure in about thirty seconds, because if a server goes down, it either reroutes or it doesn't. And language models completely shatter all those assumptions. Their output varies for real, their decay builds up instead of happening straight away, and those failures are confident, plausible, and wrong while latency, error rate, and pod health stay exactly where they should be.
This book sets out the new framework for AI infrastructure. You'll find that, working through one support-assistant application across twelve chapters, everything moves from defining steady state for probabilistic systems to injecting faults at every layer, including accelerators, inference servers, model artefacts, retrieval pipelines, model APIs, autonomous agents, distributed training and the monitoring stack itself.
Just to be clear, the toolchain is intentionally kept small. The LitmusChaos orchestrates, the Chaos Mesh reaches the kernel and filesystem, and the mitmproxy handles the content-level faults that neither can touch, extended with custom injectors where nothing suitable exists. So, every experiment comes to one of these four conclusions, and then the last chapter sums them all up as a programme scorecard. Key LearningsWrite down what "working correctly" means when your system's output legitimately varies.
Break GPUs and inference servers on purpose, with a restore path that always runs. Spot the failures that keep every dashboard green and every alert silent. Check model weights for damage before loading them, not after serving starts. Catch the model release that scores better overall while getting worse where it matters. Find out whether your RAG pipeline is helping, or quietly making answers worse.
Corrupt what the model returns, then watch whether your agent notices. Stop your remediation agent from making an incident worse than it was. Prove your training checkpoints actually resume the run they claim to continue. Catch slow decay in hours instead of days by changing how you alert. Table of ContentWhy AI Systems Fail Differently?Steady State for Probabilistic SystemLab and the ToolchainAccelerator and Compute ChaosInference Serving ChaosModel Artifacts, Registries, and Version SkewRetrieval and Context ChaosModel-Boundary ChaosAgentic and Multi-Agent ChaosDistributed Training and Choas Fine-TuningObserving Probabilistic DegradationOperationalizing Chaos