Applications Systems
SAM Windows Service Monitor Does Not Detect Brief Service Restarts or Fire Alerts
The SAM Windows Service Monitor does not record status change events or trigger alerts for a Windows service that restarts briefly, even though the service restart is confirmed in the Windows Event Viewer on the monitored node. This occurs when the SAM component monitor polling interval is longer than the actual service downtime window, causing all polling cycles to miss the brief Down state entirely.
First published date
Last published date
Overview
Symptoms
-
The SAM Application Details page shows no application events for the monitored service, despite the service restarting on a regular schedule (e.g., nightly).
-
Active Application Alerts for the application shows 0, and no alert notifications are sent.
-
The Windows Event Viewer on the monitored node confirms the service is stopping and starting as expected (e.g., stops at 3:00 AM, starts at 3:02 AM — approximately 2 minutes of downtime).
-
The alert associated with the service monitor is enabled, correctly scoped, has no sustain time, no time-of-day restriction, and evaluates every 60 seconds — but it never fires.
-
Other alerts on the SolarWinds Platform are functioning normally.
-
Poll Now from the Application Details page returns a successful result, confirming credentials and connectivity are not the issue.
Relevant Logs and Diagnostic Data
The following logs and data sources are relevant when troubleshooting this issue:
From Agent Diagnostics (if agent-polled)
|
Log / Path |
Purpose |
|---|---|
|
APM\ApplicationLogs\AppIdXX\ |
SAM application debug logs — shows what the agent sees when polling the service (requires Debug mode enabled on the application monitor) |
|
Agent\SolarWinds.Agent.Service.exe.log |
Agent health — job execution, GetDetails responses, plugin startup |
|
JobEngine\SolarWinds.JobEngineService_v2.log |
Agent-side Job Engine — infraPing heartbeats, job scheduling |
From Poller Diagnostics (APE that polls the node)
|
Log / Path |
Purpose |
|---|---|
|
LogFiles\APM\APM.BusinessLayer.log |
SAM Business Layer — processes poll results received from the agent, handles status calculations |
|
LogFiles\APM\APM.BusinessLayer.ThresholdCalculation.log |
Threshold calculation details — how SAM evaluated component thresholds |
From MPE Diagnostics (Main Polling Engine)
|
Log / Path |
Purpose |
|---|---|
|
LogFiles\Alerting.Service.V2.log |
Alert condition evaluation — did the alert engine evaluate the trigger condition (requires DEBUG via Log Adjuster) |
|
LogFiles\ActionsExecutionAlert.log |
Alert action execution — did it attempt to send an email or run an action (requires DEBUG via Log Adjuster) |
|
LogFiles\Core.BusinessLayer.log |
Business Layer processing — handles email dispatch and action processing |
From Node Diagnostics (CSV exports)
|
CSV File |
Purpose |
|---|---|
|
APM_ComponentStatus_CS_Detail_XXXX.csv |
Component-level polling history — timestamps, availability, percent availability |
|
APM_ApplicationStatus_Detail_XXXX.csv |
Application-level polling history |
|
APM_Application_XXXX.csv |
Application configuration — Unmanaged status (Maintenance Mode check) |
|
APM_CurrentComponentStatus_XXXX.csv |
Current component status — polling protocol, error codes |
|
APM_Component_XXXX.csv |
Component configuration — component name, application ID, template ID |
|
APM_ComponentDefinition_XXXX.csv |
Component definition — component type, description |
|
Pollers_XXXX.csv |
Node polling configuration — polling method (Agent/ICMP/WMI), enabled pollers |
From MPE Diagnostics (CSV exports)
|
CSV File |
Purpose |
|---|---|
|
AlertConfigurations.csv |
Full alert definitions — trigger conditions, sustain time, scope, evaluation frequency, reset conditions |
|
AlertHistoryView.csv |
Alert trigger history — confirms whether the alert has fired and when |
|
Events.csv |
Platform event history — confirms whether SAM recorded any status change events (EventType 504–510) for the node |
Product section
Cause
The default SAM component monitor polling interval is 300 seconds (5 minutes). If the monitored service restarts and recovers in less time than the polling interval (e.g., 2 minutes), it is possible for every polling cycle to occur while the service is running. SAM never observes the service in a Down state, no status change is recorded in the database, and the alert engine has no condition to trigger on.
Example timeline with a 5-minute polling interval and a 2-minute service restart:
|
Time |
Event |
|---|---|
|
~2:57 AM |
SAM polls the service → status is Up |
|
3:00 AM |
Service stops (confirmed in Windows Event Viewer) |
|
3:02 AM |
Service starts (confirmed in Windows Event Viewer) |
|
~3:02 AM |
SAM polls the service → status is already Up |
In this scenario, the entire downtime window falls between two consecutive polling cycles. SAM records Availability = 1 (Up) on every poll, and the alert trigger condition (e.g., Status != 1) is never met.
This behavior may appear suddenly if the service's restart duration shortens (e.g., from 5+ minutes to under 2 minutes) due to changes in the service configuration, hardware performance, or the scheduled task that initiates the restart.
Resolution
Verification Steps
Before applying the resolution, the following data points can be used to confirm this is the root cause:
1. Confirm the polling interval
Check the SAM component polling timestamps. Navigate to the Application Details page and review the component status history, or query the database:
-- Scripts are not supported under any SolarWinds support program or service.
-- Scripts are provided AS IS without warranty of any kind. SolarWinds further
-- disclaims all warranties including, without limitation, any implied warranties
-- of merchantability or of fitness for a particular purpose. The risk arising
-- out of the use or performance of the scripts and documentation stays with you.
-- In no event shall SolarWinds or anyone else involved in the creation,
-- production, or delivery of the scripts be liable for any damages whatsoever
-- (including, without limitation, damages for loss of business profits, business
-- interruption, loss of business information, or other pecuniary loss) arising
-- out of the use of or inability to use the scripts or documentation.
SELECT TOP 50 ComponentID, Availability, PercentAvailability, Timestamp
FROM APM_ComponentStatus_CS_Detail
WHERE ComponentID = <YourComponentID>
ORDER BY Timestamp DESC
If the timestamps are spaced exactly 5 minutes apart (e.g., :02, :07, :12, :17) and all show Availability = 1, the polling interval is 300 seconds and no Down state has been recorded.
2. Confirm the service downtime duration
On the monitored node, open Windows Event Viewer and locate the service stop and start events. Calculate the actual downtime window and compare it to the polling interval.
3. Confirm the alert configuration
Navigate to Alerts & Activity > Manage Alerts and review the alert associated with the service monitor:
|
Setting |
Expected Value |
|---|---|
|
Enabled |
True |
|
Trigger Condition |
Status != 1 (or equivalent) |
|
Sustain Time |
None (immediate) or shorter than the downtime window |
|
Time-of-Day Schedule |
None, or includes the time when the service restarts |
|
Evaluation Frequency |
60 seconds (default) |
If all of these settings are correct, the alert is not the issue — it simply never receives a status change to evaluate.
4. Confirm the alert history
Navigate to Alerts & Activity > Alert Summary or query the database:
-- Scripts are not supported under any SolarWinds support program or service.
-- Scripts are provided AS IS without warranty of any kind. SolarWinds further
-- disclaims all warranties including, without limitation, any implied warranties
-- of merchantability or of fitness for a particular purpose. The risk arising
-- out of the use or performance of the scripts and documentation stays with you.
-- In no event shall SolarWinds or anyone else involved in the creation,
-- production, or delivery of the scripts be liable for any damages whatsoever
-- (including, without limitation, damages for loss of business profits, business
-- interruption, loss of business information, or other pecuniary loss) arising
-- out of the use of or inability to use the scripts or documentation.
SELECT TOP 50 AlertHistoryID, EventType, Message, TimeStamp, AlertActiveID, AlertObjectID, ActionID
FROM AlertHistory
WHERE AlertObjectID = <YourAlertID>
ORDER BY TimeStamp DESC
If 0 entries are returned, the alert has never fired because SAM has never recorded a non-Up status for the application.
5. Confirm the alerting engine is healthy
Review the AlertHistoryView in the diagnostics or the Alert Summary page. If other alerts are actively firing, the Alerting Service is functioning normally and the issue is isolated to this specific monitor.
Resolution
Option 1: Reduce the SAM Component Polling Interval (Recommended)
Reduce the polling interval for the Windows Service Monitor component so that at least one polling cycle falls within the service downtime window.
-
Navigate to the Application Details page for the affected application.
-
Click Edit Application Monitor.
-
Select the Windows Service Monitor component.
-
Under the component settings, change the polling interval from 300 seconds to a value shorter than the service downtime window. For a 2-minute restart, a polling interval of 60 seconds is recommended.
-
Click Submit to save.
Performance Note: Reducing the polling interval for a single component has minimal impact on overall polling engine performance. However, if applying this change across many components or application monitor templates, monitor the polling engine utilization via Settings > My Deployment > Polling Engines to ensure it remains within the supported capacity.