AgentOps Is the New Day-2 Operations: A Control Model for AI Agents
Explains AgentOps as the new Day-2 operations discipline for managing AI agents in production, covering lifecycle, observability, governance, and incident response.
Paul Bryant is a Principal Multicloud Architect at Dell Technologies, sharing expert insights on hybrid cloud, HCI, and enterprise IT transformation.
176 articles from this blog
Explains AgentOps as the new Day-2 operations discipline for managing AI agents in production, covering lifecycle, observability, governance, and incident response.
Guide to configuring multi-tenant GPU scheduling on Kubernetes using NVIDIA Run:ai, covering quotas, fairshare, priority, and preemption.
Explains how to combine MCP and A2A protocols in a layered enterprise agent architecture for scalable agent-tool and agent-agent collaboration.
Guide to tuning NVIDIA Triton Inference Server for throughput and latency using Perf Analyzer and Model Analyzer.
Guide to adding NVIDIA NeMo Guardrails to a production LLM endpoint, covering architecture, enforcement, and operational controls.
Step-by-step guide to deploying NVIDIA vGPU on VMware vSphere, including validation and configuration steps.
Guide to monitoring NVIDIA GPUs using DCGM Exporter, Prometheus, and Grafana for production Kubernetes clusters.
Explains why enterprise AI agents need a control plane with identity, policy, and observability beyond just prompts.
A framework for evidence-based workload placement across cloud, VCF, Azure Local, Nutanix, and bare metal, avoiding religious debates.
Guide to deploying VMware NSX on NVIDIA Spectrum fabric, covering underlay setup, BGP routing, and validation for NSX 4.x.
Analysis of specification gaming in agentic AI where benchmark cheating becomes a production security breach, using recent OpenAI and Hugging Face incidents.
Designing a governed forensic AI platform for cyber defense that balances permissive evidence analysis with strict security controls.
Explains why agent observability requires trajectory-level detection of autonomous system divergence, not just traditional logging.
Analysis of OpenAI's cyber incident, arguing it was a control-system failure, not a rogue AI, with lessons for enterprise agentic system design.
Explains why AI agents with shell access and credentials must be secured using zero-trust architecture, treating them as privileged insiders.
Analysis of GPU multi-tenancy security controls including MIG, vGPU, confidential computing, and isolation strategies for AI platforms.
Guide to building per-tenant GPU telemetry, showback, and capacity evidence for AIaaS platforms.
A troubleshooting guide for diagnosing why GPUs remain idle during enterprise inference, focusing on systematic root-cause analysis.
Explains how MCP and A2A protocols fit into an enterprise agent architecture, with MCP for tool access and A2A for agent-to-agent delegation.
A tutorial on designing reliable AI agent tools with clear naming, schemas, error handling, and evaluation.