Skip to content

HttpClient connection pool missing idle timeout #1290

Description

@4lexBaum

Summary

On BTP CF (US30, GCP), we observed 1000+ requests in the last 24 hours taking ~120 seconds across multiple CAP Java services. Each request eventually succeeds after an automatic retry, but the ~120s stall causes significant user-visible latency.

Symptom

A CAP Java service calls an external service via OData. The first attempt fails after ~120 seconds with:

{
  "msg": "I/O exception (java.net.SocketException) caught when processing request to {s}->https://<masked>.us30.<masked>.cloud.sap:443: Connection timed out",
  "logger": "org.apache.http.impl.execchain.RetryExec",
  "level": "INFO",
  "written_at": "2026-09-29T15:14:49.995Z"
}

RetryExec then immediately retries with a fresh connection, which succeeds. The target service logs show no trace of the first attempt — the request never arrived.

This pattern is observed across 7–8 different service pairs, all running CAP Java on CF/BTP, all showing the same ~120s delay fingerprint.

Current Assumption (to be confirmed by Cloud SDK team)

This is our working assumption based on log analysis and Cloud SDK Java source code inspection

Network layer hypothesis

CF/BTP egress traffic passes through a NAT gateway. NAT gateways have an idle connection timeout — the exact value likely differs between hyperscalers (AWS, Azure, GCP) and is not publicly documented for BTP. When a pooled TCP connection sits idle past this threshold, the NAT silently drops its flow table entry without sending a TCP RST or FIN. The socket remains in ESTABLISHED state on the Java side (half-open socket).

When RetryExec tries to reuse this stale socket, it calls socket.write() to send the HTTP request. Java's SO_TIMEOUT only applies to socket.read() — there is no application-level timeout on write(). The OS TCP stack retransmits until it exhausts its retry budget (on the observed Diego cells), then raises ETIMEDOUT → SocketException: Connection timed out.

Cloud SDK source — missing idle timeout configuration

Inspecting the Cloud SDK Java source (github.com/SAP/cloud-sdk-java) shows that DefaultApacheHttpClient5Factory only sets a single 2-minute overall timeout, applied to connectTimeout and socketTimeout. The available ConnectionConfig methods for controlling connection pool lifetime are not used:

// DefaultApacheHttpClient5Factory.java — current state
PoolingHttpClientConnectionManagerBuilder.create()
    .setDefaultConnectionConfig(ConnectionConfig.custom()
        .setConnectTimeout(timeout)   // 2 minutes
        .setSocketTimeout(timeout)    // 2 minutes
        // setIdleTimeout(...)             ← not set
        // setTimeToLive(...)              ← not set
        // setValidateAfterInactivity(...) ← not set
        .build())
    ...
    .build();

Apache HttpClient's default when these are not set is no idle timeout — connections in the pool live indefinitely. Combined with the HttpClient cache TTL of 1 hour (expireAfterAccess), the same connection pool (and its potentially stale connections) is reused for up to 1 hour.

The 2-minute overall timeout is larger than the observed ~120s OS TCP timeout, meaning the OS always fires first and the application-level timeout never has a chance to act.

The same gap exists in DefaultHttpClientFactory (HC4).

Questions

  1. Can you confirm that DefaultApacheHttpClient5Factory and DefaultHttpClientFactory (Cloud SDK Java) intentionally do not set setIdleTimeout / setTimeToLive / setValidateAfterInactivity on the connection pool?
  2. Do you know the NAT idle timeout values for CF/BTP on GCP, AWS and Azure? This would help determine the right default value.
  3. Can Cloud SDK either:
    • Set a safe default idle timeout that works across all hyperscalers, or
    • Expose a configuration property (e.g. via application.yaml) so applications can set it without overriding internal SDK internals?

A custom workaround at the application level (overriding HttpClientFactory) appears technically possible but has unclear side effects given the complexity of the SDK's connection management — a proper fix at the Cloud SDK level is needed.

References

  • Apache HttpClient 5 connection management docs: https://hc.apache.org/httpcomponents-client-5.6.x/connection-management.html
  • Cloud SDK Java source DefaultApacheHttpClient5Factory.java: cloudplatform/connectivity-apache-httpclient5/src/main/java/com/sap/cloud/sdk/cloudplatform/connectivity/DefaultApacheHttpClient5Factory.java
  • Cloud SDK Java source DefaultHttpClientFactory.java: cloudplatform/connectivity-apache-httpclient4/src/main/java/com/sap/cloud/sdk/cloudplatform/connectivity/DefaultHttpClientFactory.java

Steps to Reproduce

Since it does not happen on every single request, there is no way for us to reproduce it.

Expected Behavior

Cloud SDK could either:

  • Set a safe default idle timeout that works across all hyperscalers, or
  • Expose a configuration property (e.g. via application.yaml) so applications can set it without overriding internal SDK internals?

Used Versions

  • Java and Maven version via mvn --version: 21.0.+
  • SAP Cloud SDK version: >5.31
  • CAP version: 4.9.4

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions