Network Management
Long delays when remanaging nodes on an Additional Polling Engine
In SolarWinds Observability Self-Hosted, remanage actions can take a long time to take effect when an Additional Polling Engine (APE) is over its effective licensed-entity limit. The delay is caused by prolonged EntityWatcher processing while the collector sorts entities, particularly ICMP-only nodes. Rebalancing ICMP-only nodes across polling engines with available headroom can restore normal remanage behavior.
First published date
Last published date
Overview
A remanage request may be saved successfully but not acted on until the affected polling engine completes its next EntityWatcher refresh cycle. Under normal conditions, EntityWatcher completes its pass within seconds and runs approximately every 30 seconds. In the affected environment, a refresh cycle took more than one hour, so nodes remained unmanaged or in an unknown state for an extended period.
This condition can also delay other operations that depend on timely polling-state updates, including NCM jobs and the transition of a node to Down status after it becomes unavailable.
Symptoms
One or more of the following symptoms may occur:
-
Remanaging a node completes slowly or appears not to complete.
-
A node remains Unmanaged or Unknown after the remanage action.
-
Polling-state changes are delayed.
-
NCM jobs that include nodes managed by the affected APE take a long time.
-
The issue affects some polling engines but not others.
-
Restarting the affected services, restarting the server, or running the Configuration Wizard does not resolve the behavior.
Product section
Cause
Each polling engine runs EntityWatcher as a background task to detect entity changes, including Manage and Unmanage state changes. Before applying changes, the collector can sort the entity list to protect licensed nodes while evaluating ICMP-only nodes.
ICMP-only nodes do not consume a node license. However, when an engine could be over its effective limit, the collector performs additional sorting and license-headroom checks. In the reported case, the sorting operation in SolarWinds.Orion.Core.Collector.Node.NodeEntityCreator.SortedEntities caused an EntityWatcher refresh to take approximately 1 hour and 7 minutes.
The effective limit is calculated from the engine's available license headroom plus its own licensed nodes. The following values illustrate the condition observed in the reported environment; they are not universal limits:
|
Polling-engine condition |
Total nodes |
Licensed SNMP/Agent nodes |
ICMP-only nodes |
Effective limit |
Result |
|---|---|---|---|---|---|
|
Engine with available headroom |
523 |
368 |
155 |
1,831 |
1,308 below limit |
|
Affected engine 1 |
2,939 |
1,032 |
1,907 |
2,495 |
444 over limit |
|
Affected engine 2 |
2,659 |
1,006 |
1,653 |
2,469 |
190 over limit |
|
Engine with available headroom |
1,157 |
1,121 |
36 |
2,584 |
1,427 below limit |
The affected engines were the engines over their effective limit. Moving licensed SNMP or Agent nodes alone does not resolve the condition because each moved licensed node reduces both the licensed-node count and the effective limit by the same amount.
Resolution
Resolution / Workaround
-
Identify the polling engines that are over their effective licensed-entity limit and the engines that have sufficient headroom.
-
Rebalance ICMP-only nodes away from the affected engines and onto engines with available headroom.
-
As a starting point, the reported environment required moving approximately 600 ICMP-only nodes from one affected engine and approximately 300 ICMP-only nodes from the other affected engine. Determine the appropriate quantity for each environment based on its node and license counts.
-
Do not rely on moving SNMP or Agent nodes as the primary corrective action; this does not reduce the over-limit condition in the required way.
-
After rebalancing, test the following from the SolarWinds Web Console:
-
Manually unmanage and remanage a test node on each previously affected APE.
-
Confirm that the node returns to the expected polling state promptly.
-
Run or monitor an NCM job containing a node managed by the APE.
-
-
If the issue persists, collect current diagnostics and escalate for engineering review. Include the affected engine names, node counts, license counts, timestamps for the Manage/Unmanage actions, and EntityWatcher log entries.