Jun 13Vibe with Hermes Agent — Bengaluru · RSVP
ToolsMCPBlogResearchCommunityStar on GitHub
All Tools
T
MonitoringFreeOpen Source

TELEGRAF

Plugin-driven metrics collection agent

MIT

ABOUT

ML infrastructure generates metrics from heterogeneous sources — GPU utilization from NVIDIA tools, model serving logs from HTTP endpoints, system metrics from the OS, database connection pools, and custom application metrics — all in different formats. Telegraf provides a single, plugin-based agent with 300+ input plugins (CPU, disk, nvidia-smi, Kafka, HTTP, Prometheus, MQTT, etc.) that collect metrics from every layer of the stack, transform and enrich them, and output to any monitoring backend (InfluxDB, Prometheus, Graphite, Datadog, CloudWatch). This eliminates the need for separate collection agents per source.

INTEGRATION GUIDE

1. Collect GPU utilization, memory, and temperature metrics from NVIDIA GPUs across a training cluster 2. Ingest model serving latency and throughput metrics from HTTP endpoints and message queues 3. Monitor training job resource usage (CPU, memory, disk, network) across distributed workers 4. Aggregate infrastructure health metrics from Kubernetes, Docker, and host-level sensors 5. Forward collected metrics to multiple backends simultaneously (InfluxDB for storage, Prometheus for alerting)

TAGS

monitoringmetricsobservabilitydata-collectioninfrastructure
Telegraf — AI Tool | Agentic AI For Good