Observability
Polling stops and ephemeral ports are exhausted on TCP port 17733 due to SolarWinds Collector to Job Engine v3 communication in the SolarWinds Platform
This article provides information about an issue in the SolarWinds Platform where the SolarWinds Collector service opens a very large number of TCP connections to the SolarWinds Job Engine v3 service on local port 17733, exhausting the available ephemeral (dynamic) ports on the polling engine. When this occurs, polling stops or becomes unresponsive, and other components on the same server that require new outbound connections may also begin to fail.
First published date
Last published date
Overview
On the affected polling engine, one or more of the following may be observed.
Polling and data symptoms:
- Polling completion drops to a very low value, such as 1%, or polling stops entirely.
- Node and interface data becomes stale, and interfaces may be displayed as obsolete.
- Next Poll Time does not advance for affected nodes, while Last Sync continues to update.
- Nodes are not generating Up or Down events even though the devices are reachable.
- Removal of old jobs in Job Engine v3 is extremely slow, and new jobs are only scheduled after the old jobs have finished clearing.
- Job Engine v3 retains a very high scheduled job count, for example more than 30,000 jobs.
Server and connection symptoms:
- A very high number of ESTABLISHED TCP connections from SolarWinds.Collector.Service.exe to 127.0.0.1:17733, in the thousands or tens of thousands.
- The SolarWinds Platform Web Console becomes slow, times out, or returns errors.
- A server restart or a full restart of all SolarWinds services temporarily restores polling, but the symptom returns.
- Winsock or socket buffer errors appear for unrelated components on the same server.
Common triggers:
- A High Availability pool failover, in either direction, whether automatic or manually initiated. The newly active server inherits a Job Engine v3 job set that its Collector service does not recognize.
- Stopping the Collector service, deleting the Collector SQLite data store, and starting the Collector service again.
- Disabling and then re-enabling the Job Engine v3 service, which moves jobs between Job Engine v2 and Job Engine v3.
Sample log entries
Log entries below are representative examples. Timestamps, thread and correlation identifiers, hostnames, IP addresses, and entity identifiers vary by environment and are shown in generic form.
Collector service log, located at C:\ProgramData\SolarWinds\Collector\Logs\. The Collector is unable to reach the Job Engine v3 gRPC endpoint on the local port:
YYYY-MM-DD hh:mm:ss,nnn <THREAD_ID> ERROR SolarWinds.JobEngine.GrpcClient.JobEngineGrpcClientExceptionHandler - [<CORRELATION_ID>:LogGrpcClientException] The Job Engine gRPC client encountered an 'RpcException' exception - 'Status(StatusCode="Unavailable", Detail="Error starting gRPC call. HttpRequestException: An error occurred while sending the request. WebException: Unable to connect to the remote server SocketException: No connection could be made because the target machine actively refused it 127.0.0.1:17733", DebugException="System.Net.Http.HttpRequestException: An error occurred while sending the request.")'. CorrelationId: '00000000-0000-0000-0000-000000000000', [Optional]JobExecutionId: '(null)'.
Collector service log, showing that the polling plan cannot be resolved:
YYYY-MM-DD hh:mm:ss,nnn <THREAD_ID> ERROR SolarWinds.Collector.PollingController.PollingPlanResolver - Resolve polling plan failed!! We will try to resolve it in next round SolarWinds.JobEngine.Exceptions.JobEngineException: Status(StatusCode="Internal", Detail="Error starting gRPC call. HttpRequestException: An error occurred while sending the request. WebException: The underlying connection was closed: The connection was closed unexpectedly.", DebugException="System.Net.Http.HttpRequestException: An error occurred while sending the request.")
Job Engine v3 service log, located at C:\ProgramData\SolarWinds\Logs\Orion\. Requests are aborted or arrive too slowly while the port range is exhausted:
YYYY-MM-DD hh:mm:ss,nnn .NET TP Worker ERROR Microsoft.AspNetCore.Server.Kestrel - Unhandled exception while processing <CONNECTION_ID>. Grpc.Core.RpcException: Status(StatusCode="Unavailable", Detail="Error starting gRPC call. HttpRequestException: An error occurred while sending the request. IOException: The request was aborted. IOException: Unable to read data from the transport connection: An existing connection was forcibly closed by the remote host.. SocketException: An existing connection was forcibly closed by the remote host.", DebugException="System.Net.Http.HttpRequestException: An error occurred while sending the request.")
Certificate or gRPC service method errors may also be logged while the connection backlog is present:
YYYY-MM-DD hh:mm:ss,nnn .NET TP Worker ERROR Grpc.AspNetCore.Server.ServerCallHandler - (null) Error when executing service method 'GetCertificateThumbprint'. System.IO.IOException: The request stream was aborted. ---> Microsoft.AspNetCore.Connections.ConnectionAbortedException: The HTTP/2 connection faulted. --- End of inner exception stack trace ---
Once ephemeral ports are fully consumed, other SolarWinds components on the same server report Winsock resource errors. This entry may be seen in the Collector or other platform service logs:
YYYY-MM-DD hh:mm:ss,nnn <THREAD_ID> FATAL SolarWinds.Orion.Core.Common.RemoteFeatureManager - Unable to load Feature Manager data from SWIS System.InsufficientMemoryException: Insufficient winsock resources available to complete socket connection initiation. ---> System.Net.Sockets.SocketException: An operation on a socket could not be performed because the system lacked sufficient buffer space or because a queue was full <IP_ADDRESS>:17777
Repetitive entity warnings may also fill the Collector log and make other entries harder to locate. These entries are noise and are not the cause of this issue:
YYYY-MM-DD hh:mm:ss,nnn [EntityWatcher] ERROR SolarWinds.Collector.BusinessLayer.EntityWatcher - Parent entity <ENTITY_TYPE>:<ID> is missing for entity <ENTITY_TYPE>:<ID>
Confirming whether this issue is present
Run the following command on the affected polling engine while the symptom is active. It returns the number of connections on port 17733:
netstat -ano | find "17733" | find /c ":17733"
Interpret the result as follows:
- A result in the thousands or tens of thousands indicates this issue.
- A result in the tens, for example between 8 and 50, means the ephemeral port range is not being consumed on port 17733 and this article does not apply. Continue troubleshooting the reported symptom as a separate issue.
To review the overall ephemeral port utilization on the server, use:
netstat -ano | find /c "TCP"Product section
Cause
The SolarWinds Collector service communicates with the SolarWinds Job Engine v3 service over local TCP port 17733 using gRPC-web over HTTP/1.1. Because this transport does not support HTTP/2 connection multiplexing, every request requires a separate TCP/IP connection and consumes one ephemeral port.
When the Job Engine v3 service contains a large number of scheduled jobs that are not present in the Collector service, the Job Engine v3 service returns the results of those jobs to the Collector. The Collector does not recognize the jobs, so it issues an individual job removal request back to the Job Engine v3 service for each one, without any rate limiting or batching. Each of those removal requests opens a new connection on port 17733.
The legacy SolarWinds Job Engine v2 service uses a transport that pools connections and applies request throttling, which prevented this behavior in earlier architectures. The Job Engine v3 service had no equivalent throttling, so a large volume of job removal requests could consume the entire ephemeral port range on the server.
This mismatch between the Collector job list and the Job Engine v3 job list most commonly occurs after a High Availability pool failover, after the Collector SQLite data store is deleted, or when the Job Engine v3 service is disabled and re-enabled.
A separate certificate validation defect in the SolarWinds Certificate Management Service could produce the same high connection count on port 17733 by leaving requests stuck at server or client certificate validation. That defect was addressed in SolarWinds Platform 2026.2 and is distinct from the job removal behavior described above.
Affected versions include SolarWinds Platform 2026.1, 2026.1.1, 2026.2, and 2026.2.1. The certificate validation improvement shipped in 2026.2, and 2026.2.1 contained only the certificate fix, so neither release fully addresses the job removal behavior described in this article.
Resolution
Permanent resolution
Upgrade the affected polling engines to SolarWinds Platform 2026.2.2 or later.
In 2026.2.2, job removal requests from the Collector service to the Job Engine services are batched rather than sent individually. Job removal identifiers are queued per Job Engine version and flushed in capped batches, or after a short timer interval, whichever occurs first. This reduces the number of connections opened on port 17733 by approximately two orders of magnitude for the same volume of job removals, and prevents the ephemeral port range from being consumed.
The published fix description for this release is: communication between the collector and the job engine on port 17733 no longer leads to port exhaustion.
Note that the upgrade must be to 2026.2.2 or later. Upgrading only to 2026.2 or 2026.2.1 does not resolve this issue.
Upgrade and verification steps:
- Upgrade your SolarWinds Platform environment to version 2026.2.2 or later.
- Stop the SolarWinds Collector service from the SolarWinds Platform Service Manager.
- Navigate to C:\ProgramData\SolarWinds\Collector\Data and delete the contents of that folder.
- Start the SolarWinds Collector service. The data store is rebuilt automatically and scheduled jobs are restored from the SolarWinds Platform database.
- Allow the environment to run for approximately one hour, then confirm the following.
Verification after upgrade:
- The connection count on port 17733 remains in the tens during peak activity, not in the thousands. Use netstat -ano | find "17733" | find /c ":17733".
- Old jobs are removed from the Job Engine v3 service and the scheduled job count decreases to an expected level.
- Polling completion returns to a normal value and node data updates on schedule.
- The Collector and Job Engine v3 logs no longer contain the connection errors listed in the Overview section. Warnings stating that a job was not found in the queue or in the database are expected during job cleanup and are not an indication of a problem.
Interim workaround
If the deployment cannot be upgraded to 2026.2.2 immediately, the Job Engine v3 service can be disabled so that all polling jobs run on the Job Engine v2 service. This is intended as a temporary measure until the upgrade is completed.
- Open Advanced Configuration settings on the SolarWinds Platform server.
- Navigate to SolarWinds.JobEngine.Service.v3.Settings.
- Set EnableJobEngineV3 to disabled.
- Restart the SolarWinds services on the affected polling engines.
Behavior while the Job Engine v3 service is disabled:
- All polling jobs run on the Job Engine v2 service, as in earlier versions.
- All product functionality and features continue to operate.
- There are no expected operational limitations or side effects.
- High Availability is not affected. The Job Engine v3 service continues to run, but no jobs are scheduled to it.
- The Collector service communicates only with the Job Engine v2 service.
Re-enable EnableJobEngineV3 after the upgrade to 2026.2.2 or later has been completed.
Because polling is progressively being migrated to the Job Engine v3 service in future releases, disabling it is not a long-term solution and the upgrade should still be planned.
Recovering a polling engine that is currently exhausted
Use these steps to restore polling when the ephemeral port range is already consumed and the deployment has not yet been upgraded.
- Stop the following services from the SolarWinds Platform Service Manager, in this order: SolarWinds Collector service, SolarWinds Job Engine v2 service, SolarWinds Job Engine v3 service.
- Delete the contents of C:\ProgramData\SolarWinds\Collector\Data.
- Delete the Job Engine v2 data files from C:\ProgramData\SolarWinds\JobEngine.v2\Data.
- Delete JobEngine.db from C:\ProgramData\SolarWinds\JobEngine.v3\Data.
- Start the SolarWinds Certificate Management Service and confirm it is running.
- Start the SolarWinds Job Engine v3 service, then the SolarWinds Job Engine v2 service, then the SolarWinds Collector service.
Deleting these data files is safe. The Job Engine and Collector services recreate the databases and restore scheduled jobs from the SolarWinds Platform database on the next startup. Allow up to approximately one minute after the services start for the files to be recreated.
Important: on versions earlier than 2026.2.2, stopping the Collector service, deleting the Collector data store, and starting the Collector service is also a trigger for this issue. On an affected version this procedure may restore polling temporarily and then cause the connection count to rise again. Treat it as recovery only, and plan the upgrade to 2026.2.2 or later as the resolution.
If the environment uses High Availability and cannot be upgraded before the next planned failover, clearing the Collector, Job Engine v2, and Job Engine v3 data files on the standby server before the failover may reduce the likelihood of the issue occurring, because the incoming active server starts with a clean job state.
Additional environment consideration
Expanding the Windows dynamic port range provides additional headroom on a busy polling engine, but it does not prevent this issue and is not a substitute for the upgrade.
netsh int ipv4 set dynamicport tcp start=40000 num=25536
For more information about ephemeral port utilization on SolarWinds Platform servers, see the article on ephemeral port exhaustion on SolarWinds Platform servers.