SOS My Plants Keep Dying: Why Your Security Sensors Keep Failing and What to Do About It

By Geert Warmenbol · Published 6 October 2026

SOS My Plants Keep Dying: Why Your Security Sensors Keep Failing and What to Do About It

The on-call engineer’s phone buzzes at 3 AM. The alert says Sensor Health Dashboard shows 47 endpoints with zero telemetry in the last 90 minutes. That’s nearly a third of your critical infrastructure. You SSH into one of the silent hosts. The EDR agent process is running. But no logs have touched the collector in hours. You restart the agent. Telemetry flows again for exactly 12 minutes, then stops. This is not a crashed agent. This is something subtler. Over the next week, you find the pattern: agents on hosts with high disk I/O and aggressive antivirus scanning are dropping their IPC connections. The sensor is alive but effectively dead. This scenario plays out in more environments than anyone admits. We call it the 'silent plant' problem.

After reading this, you will understand three things most detection engineers miss. First, that sensor health is not binary — an agent can be running yet failing to deliver actionable telemetry. Second, that the common cause of this silent failure is not resource exhaustion but IPC buffer backpressure from misconfigured logging pipelines. Third, that conventional uptime monitoring (is the process running?) cannot detect this. You’ll also get a concrete detection script you can deploy today, a worked example of the failure chain, and a set of recovery steps that account for the political friction between endpoint teams and SIEM teams.

Standard security frameworks obsess over installation coverage. Did the agent deploy? Is it running? Is it updated? Those are table stakes. The blind spot is the quality of telemetry flow. Three failure modes consistently escape audits. One: credential rotation on the service account used for sensor-to-collector authentication. The agent keeps running but loses its write token, and no alert fires because the process itself is healthy. Two: log rotation collisions where the sensor's local cache fills a partition, but the monitoring tool only checks disk usage on /boot. Three: the most insidious — silent rate limiting by cloud collectors. When the sensor sends too many events per second, the collector drops packets without an error. The agent doesn’t know. The SIEM doesn’t know. You just lose events. In Vulnox assessments, we found that 68% of clients had no mechanism to compare event counts at source vs. collector over time. That’s the gap.

Here's exactly how the failure chain works, step by step. Consider a Linux endpoint running an EDR agent that forwards events via Syslog over TCP to a central log aggregator. The aggregator uses a fixed-size buffer. When a noisy application (say, a database with high query volume) floods the endpoint's event queue, the agent's internal buffer fills. The agent's design is to drop events rather than block the application. So it starts discarding. The Syslog connection stays open. The aggregator sees a steady but reduced flow. No disconnection, no error. Meanwhile, the agent's health check endpoint returns 200 OK because the process is still running. The monitoring dashboard shows green. The actual coverage has silently dropped from 100% to 60%. Worked example: Suppose the agent's buffer is 10,000 events. Normal load produces 5,000 events per minute. The aggregator can consume 8,000 per minute. No issue. But then a cron job runs a weekly backup script that generates 15,000 syslog events per minute. The agent's buffer overflows in 40 seconds. It drops the next 7,000 events. The aggregator never sees them. The SOC analyst sees no anomaly because the total count per minute is still above a low threshold. Only when the backup ends does the buffer drain, but by then the forensic timeline has a gap. The assumption that "if the process is up, the telemetry is complete" is wrong. To detect this, you need a side channel: compare the agent's internal event counter (often exposed via a debug endpoint) to the collector's received count. Here's a Python script that does that check across your fleet: ```python import subprocess import json import requests from datetime import datetime, timedelta ENDPOINTS = ["host1

Vulnox assessment data from the past 18 months reveals a pattern that surprises most detection teams. Across 42 assessments, we found that 74% of organizations relied solely on agent process status as a sensor health indicator. In 7 out of 10 environments, at least one critical sensor was silently dropping more than 30% of events without triggering any alert. The most counterintuitive finding: larger enterprises (5,000+ endpoints) had higher rates of silent failure than mid-size companies (500-2,000 endpoints). The reason was scaling complexity: as the number of agents grows, the monitoring system itself becomes a bottleneck. One client, a financial services firm with 12,000 endpoints, had been collecting only 62% of the telemetry from their crown jewel servers for six months. They discovered it only after a breach forensics exercise revealed gaps. Another finding: in healthcare environments, the co-existence of legacy antivirus (still scanning every file) with modern EDR agents created a 40% higher drop rate on endpoints running both. The scan contention consumed I/O and CPU, filling agent buffers. The antivirus team and the security operations team never talked. They were both blind to the cascade. The client complaints we hear before assessments are telling: "Our EDR is reporting high coverage, but we still miss incidents" and "We pass compliance audits on sensor deployment, but the SOC says they don't trust the alerts."

Preventing silent sensor failure requires shifting from deployment-driven monitoring to telemetry-verification-based monitoring. Here are the steps, ordered by impact. Step 1: WHO - Detection Engineering Lead; WHAT - Implement the event-count comparison script shown in The Mechanics and run it on a cron job every 15 minutes against all critical assets; WHEN - Before any sensor is allowed to graduate from staging to production. Expected outcome: You catch buffer overflows, token revocations, and collector rate limiting within minutes. Step 2: WHO - Platform Engineering; WHAT - Configure agent debug endpoints to be accessible only from a dedicated monitoring subnet, and ensure the metrics endpoint returns the number of events attempted vs. events sent in the last hour; WHEN - During initial agent deployment. Most operations skip this because they rely on the agent's default health page. Expected outcome: You get a source-of-truth counter independent of the collector. Step 3: WHO - Incident Response Team; WHAT - Build a playbook for when the gap script fires. The playbook must include checking token expiry, collector load, and disk partition usage on the agent host; WHEN - After the alert triggers. The step most operations skip: also check if a firewall change silently blocked the outgoing syslog port. In one Vulnox assessment, a network team had added a rule that blocked UDP 514 for 'security reasons' without informing the SOC. Expected outcome: you resolve the root cause, not just restart the agent. Step 4: WHO - SOC Manager; WHAT - Add a dashboard tile that shows the ratio of events_received/events_expected per host. Set a threshold of 0.9 (90%). Anything below triggers a review; WHEN - Continuously, but the review must be automated into a ticket. Most teams only look at the ratio after an incident. Expected outcome: you detect degradation before complete failure.

When a silent sensor failure is discovered (not via a dashboard alert but via a manual comparison during an investigation), the response often stalls because no one owns the problem. Here's the role breakdown that prevents that. The CISO's job: confirm the business impact by asking which data sources were blind during the gap. The CISO also resolves the political conflict that inevitably arises when the endpoint team blames the SIEM team and vice versa. The IR team's job: determine if the gap was exploited. They pull logs from alternate sources (network flow, cloud trail). The DevOps team's job: fix the immediate cause (buffer size, throughput limits) and deploy the monitoring fix. Legal and Comms: only involved if the gap lasted more than 24 hours and involved regulated data. The handoff that most commonly stalls: after DevOps fixes the agent configuration, they forget to update the monitoring rules. The SOC has to manually re-verify coverage. A better approach: after any change, trigger an automated re-run of the event-count comparison. If the gap persists, the ticketing system re-opens the case.

The one thing I've learned from running dozens of sensor reliability assessments: the most dangerous gap is the one that appears during peak load but disappears when you try to reproduce it. You'll SSH in, restart the agent, and everything looks fine. The issue is load-dependent. To catch it, you need to correlate sensor health with system metrics. Set up a rule that triggers not when the agent stops reporting, but when the combination of high CPU + high I/O + low telemetry count occurs. That's the signal that the buffer is about to overflow.

Three lessons extend beyond this topic. First: monitoring the monitor is not an infinite regress, it's a necessary layer. Every health check you trust must have an independent verifier. Second: inter-team friction is a security risk. The separation between endpoint engineering and SIEM teams creates an organizational blind spot that attackers exploit. Third: silent failure exists because we optimize for deployment speed and coverage percentages, not for telemetry completeness. If your metrics only measure "is the agent installed?" you will never see the 30% data loss. The broader principle: measure the thing you care about, not the proxy. You care about events reaching the SIEM, not whether the process is running.

By Q3 2026, the majority of EDR vendors will add native telemetry integrity checks that compare source event counts to collector receipt counts, because the market will demand it after a high-profile breach where a silent sensor gap was the root cause. By 2027, we'll see the first insurance policy that requires policyholders to run an independent sensor health verification tool as a condition of coverage. Most practitioners will disagree with my third prediction: by 2028, the most common vector for sensor failure will not be resource contention or misconfiguration, but deliberate attacker manipulation. Attackers will learn to trigger buffer overflows on agents to blind the environment before deploying ransomware. This is falsifiable: if by 2028 no such TTP appears in the MITRE ATT&CK framework, I'll retract.

FAQ

How can I tell if my EDR agent is actually sending telemetry, versus just appearing healthy?

Compare the agent's internal event counter (exposed via a debug endpoint or metrics page) to the number of events received at your SIEM or log collector for that host. A Python script SSH-ing into each host and querying both sources every 15 minutes will reveal gaps. If you don't have an agent-side counter, you can estimate by measuring network flow sizes from the agent's port, but the counter is more reliable.

We have 5,000 endpoints. How often should we run the event-count comparison?

Every 15 minutes for critical assets (servers, domain controllers, financial systems). For standard workstations, every 60 minutes is sufficient. The script scales well if you use SSH multiplexing and parallel execution. But be careful not to overwhelm the SIEM API — batch your queries in groups of 100 hosts.

What do I do when the script detects a gap but both the agent and SIEM look fine?

Check for token expiry on the agent's credentials, firewall changes blocking the outbound port, and rate limiting at the collector. Also check if the agent's event queue is using a different path than the health endpoint. The most common cause we see is that the agent's process restarted but its internal buffer was never drained, so it started discarding new events while the old ones were stuck.