dt-obs-hosts

Host and process metrics including CPU, memory, disk, network, containers, and process-level telemetry. Use when analyzing infrastructure health, resource utilization, process consumption, or host discovery. Also use when building timeseries queries for host metrics that feed into analytical workflo

By dynatrace · 1,980 installs

npx skills add dynatrace/dynatrace-for-ai --skill dt-obs-hosts

Source repository · Upstream listing

Infrastructure Hosts Skill Monitor and manage host and process infrastructure including CPU, memory, disk, network, and technology inventory. When to Use This Skill Use this skill when the user needs to: Inventory: "Show me all Linux hosts in AWS us east 1" Monitor: "What hosts have high CPU usage?" Troubleshoot: "Which processes are consuming the most memory?" Discover: "What databases are running in production?" Plan: "Track Kubernetes version distribution for upgrade planning" Cost: "Calculate infrastructure costs by cost center" Security: "Find all processes listening on port 22" Compliance: "Identify hosts running EOL Java versions" Quality: "Check data completeness for AWS hosts" Optimize: "Find rightsizing candidates based on utilization" Cross source join required: If the query must combine host data with logs or other telemetry sources (e.g. "show logs from Linux hosts with their IP addresses") → also read dt dql essentials/references/smartscape topology navigation.md before writing the query. Core Concepts Entities HOST Physical or virtual machines (cloud or on premise) PROCESS Running processes and process groups CONTAINER Kubernetes containers NETWORK INTERFACE Host network interfaces DISK Host disk volumes Metrics Categories 1. Host Metrics dt.host.cpu. , dt.host.memory. , dt.host.disk. , dt.host.net. 2. Process Metrics dt.process.cpu. , dt.process.memory. , dt.process.io. , dt.process.network. 3. Inventory OS type, cloud provider, technology stack, versions 4. Cost dt.cost.costcenter , dt.cost.product 5. Quality Metadata completeness, version compliance Alert Thresholds CPU/Memory/Disk: 80% warning, 90% critical Network: 70% high, 85% saturated Disk Latency: 20ms bottleneck Network Errors: Drop rate 1%, error rate 0.1% Swap: 30% warning, 50% critical Key Workflows 1. Host Discovery and Classification Discover hosts, classify by OS/cloud, inventory resources. OS Types: LINUX , WINDOWS , AIX , SOLARIS , ZOS → For cloud specific attributes, see [references/inventory discovery.md]( cloud specific attributes) 2. Resource Utilization Monitoring Monitor CPU, memory, disk, network across hosts. High utilization threshold: 80% warning, 90% critical Key CPU Metrics: dt.host.cpu.usage — Total CPU utilization (0 100%) dt.host.cpu.idle — CPU idle time (inverse of usage; useful for anomaly detection) dt.host.cpu.user — CPU time in user mode dt.host.cpu.system — CPU time in kernel mode dt.host.cpu.iowait — CPU waiting for I/O (Linux only) → For detailed CPU analysis, see [references/host metrics.md](references/host metrics.md cpu monitoring) → For memory breakdown, see [references/host metrics.md](references/host metrics.md memory monitoring) Disk Free Space — Find Hosts with Most/Least Free Disk 3. Process Resource Analysis Identify top resource consumers at process level. → For process I/O analysis, see [references/process monitoring.md](references/process monitoring.md process io) → For process network metrics, see [references/process monitoring.md](references/process monitoring.md process network) 4. Technology Stack Inventory Discover and track software technologies and versions. Common Technologies: Java, Node.js, Python, .NET, databases, web servers, messaging systems → For version compliance checks, see [references/inventory discovery.md](references/inventory discovery.md technology inventory) 5. Service Discovery via Ports Map listening ports to services for security and inventory. Well known ports: 80 (HTTP), 443 (HTTPS), 22 (SSH), 3306 (MySQL), 5432 (PostgreSQL) → For comprehensive port mapping, see [references/inventory discovery.md](references/inventory discovery.md port discovery) 6. Container and Kubernetes Monitoring Track container distribution and K8s workload types. Workload Types: deployment , daemonset , statefulset , job , cronjob Note: Container image names/versions NOT available in smartscape. → For K8s version tracking, see [references/container monitoring.md](references/container monitoring.md kubernetes versions) → For container lifecycle, see [references/container monitoring.md](references/container monitoring.md container inventory) 7. Cost Attribution and Chargeback Calculate infrastructure costs by cost center. → For product level cost tracking, see [references/inventory discovery.md](references/inventory discovery.md cost attribution) 8. Infrastructure Health Correlation Correlate host and process metrics for cross layer analysis. Health scoring: Critical if any resource 90%, warning if 80% → For multi resource saturation detection, see [references/host metrics.md](references/host metrics.md resource saturation) Response Construction When the user asks for data retrieval or a DQL query (e.g., "show me top hosts by CPU"), include the DQL query in the response alongside the results. Users want to see and reuse the query — it is the deliverable, not just a means to get results. When the user asks for analysis (anomaly detection, forecasting, seasonality), the analysis results are the deliverable. Focus on presenting findings clearly: Prioritize metric level findings over data collection artifacts. If an analysis tool reports data gaps alongside actual anomalies, lead with the metric behavior the user asked about and mention gaps only as supplementary context. Include host names (not just IDs) using getNodeName(dt.smartscape.host) or the get entity name tool. State the timeframe analyzed and the tools/parameters used. Analytical Workflows Host metric queries often serve as inputs to analytical tools (anomaly detection, forecasting, seasonality analysis). This skill helps construct the right DQL query; the actual analysis is performed by dedicated tools. Anomaly Detection and Pattern Analysis When users ask about "unusual behavior", "anomalies", "spikes", or "sudden changes" in host metrics, the workflow is: 1. Construct the timeseries query using this skill's patterns 2. Pass it to the appropriate analysis tool (anomaly detector, novelty detection) Choosing between detectors: adaptive anomaly detector — use when the user asks about magnitude : "spikes", "abrupt changes", "values that went above normal", "sudden jumps". It answers "did this metric cross an unexpected threshold?" and reports alert durations and peak values. timeseries novelty detection — use when the user asks about behavioral change : "unusual patterns", "something changed", "trends", "new behavior". It answers "did the shape of the signal change?" without implying a specific threshold was crossed. Response format for anomaly results: Include both the host name (resolved via getNodeName(dt.smartscape.host) or get entity name ) and the host entity ID alongside timestamps and values. Entity IDs alone are opaque to users; names alone prevent follow up queries. Novelty type selection rule: When using novelty detection, set analysisNoveltyType to only [SPIKE, CHANGE IN VALUES, TREND IN VALUES] by default. EXCLUDE GAP WITH MISSING VALUES and CHANGE IN MISSING VALUES unless the user explicitly asks about data gaps or monitoring coverage. Data gaps are infrastructure issues, not metric behavior anomalies — reporting them when the user asks about CPU or memory patterns is incorrect. Queries for analysis tools should use simple timeseries format with a single aggregated metric and appropriate time range: Avoid adding filters or field transformations that reduce the data — the analysis tools work best with complete timeseries data. Forecasting When users ask to "predict", "forecast", or "estimate future" host metrics: 1. Construct the timeseries query with sufficient historical data (e.g., 7d for short term, 30d for longer predictions) 2. Pass to the forecasting tool with the desired forecast horizon The forecast horizon (how far ahead to predict) and the historical window (how much past data the model trains on) are independent. A request like "forecast the next 2 hours" sets the horizon to 2h — it says nothing about the lookback. Always use at least 7 days of historical data regardless of how short the forecast horizon is. Too few training data points cause the forecast model to fail and fall back to raw historical values. Seasonality Detection When users ask about "seasonality", "weekly patterns", or "recurring behavior": 1. Use a longer time range (at least 14d for weekly, 30d+ for monthly) 2. Pass to the seasonal baseline anomaly detector Response format for seasonal analysis: When presenting results, include: Whether seasonal anomalies were detected (yes/no) The analysis timeframe and parameters used For each affected host: host name (not just ID), timestamps of violations, violation counts, baseline values vs actual values, and upper/lower bounds Organize results by host if multiple hosts are involved Scope Boundary — Service Level vs Host Level Metrics This skill covers host and process infrastructure metrics only . If the user asks about service level metrics (request rate, response time, error rate, service calls per minute, throughput), use dt obs services instead — even when the question involves forecasting or anomaly detection of those metrics. Redirect these to dt obs services : "service calls per minute", "request rate", "response time by service", "error rate by endpoint", "service throughput forecast". Common Query Patterns Pattern 1: Smartscape Discovery Use smartscapeNodes to discover and classify entities. Pattern 2: Timeseries Performance Use timeseries to analyze metrics over time. Pattern 3: Cross Layer Correlation Correlate host and process metrics. Pattern 4: Entity Enrichment with Lookup Enrich data with entity attributes. After lookup , reference fields with lookup. prefix. Tags and Metadata Important Notes Generic tags field is NOT populated in smartscape queries Use specific tag fields: tags:azure[ ] , tags:environment Use custom metadata: host.custom.metadata[ ] Available Tags Azure Tags: tags:azure[dt owner team] , tags:azure[dt cloudcost capability] Environment: tags:environment Custom Metadata: host.custom.metadata[OperatorVersion] , host.custom.metadata[Cluster] Cost: dt.cost.costcenter , dt.cost.product → For complete tag reference, see [references/inventory discovery.md]( tags and metadata) Cloud Specific Attributes AWS cloud.provider == "aws" aws.region , aws.availability zone , aws.account.id aws.resource.id , aws.resource.name aws.state (running, stopped, terminated) Azure cloud.provider == "azure" azure.location , azure.subscription , azure.resource.group azure.status , azure.provisioning state azure.resource.sku.name (VM size) Kubernetes k8s.cluster.name , k8s.cluster.uid k8s.namespace.name , k8s.node.name , k8s.pod.name k8s.workload.name , k8s.workload.kind → For multi cloud analysis, see [references/inventory discovery.md](references/inventory discovery.md multi cloud hosts) Best Practices 1. Use percentiles (p95, p99) for latency; max() for limits; avg() for trends 2. Set multi level thresholds (warning 80%, critical 90%) 3. Filter early in the pipeline; limit results with limit N 4. Aggregate before enrichment (lookup) 5. Use getNodeName(dt.smartscape.host) for human readable host names; getNodeName(dt.smartscape.process) for processes 6. Convert bytes to GB: / 1024 / 1024 / 1024 ; round with round(value, decimals: 1) Time windows: Real time: 5 15 min Trends: 1 7 days Capacity planning: 30 90 days Limitatio