Skip to content

Major performance regression in client-v2 #2516

Description

@alekkol

Summary

After migrating from the ClickHouse Java client v1 to v2 we observed a major performance regression: more than 2 times less throughput. We use ClickHouse as a pre-aggregation layer, and run analytical queries that may return 10^6-10^9 rows for later processing. Because our pipeline depends on low-latency, low-overhead reads, even small regressions translate into major throughput losses.

Reproduction

The benchmark suite in this repository measures the general query performance differences between v1 and v2 but does not cover the most performance-sensitive scenario: retrieving column values using strongly-typed getters (e.g., the reason java.sql.ResultSet#getLong() exists). In such cases, the new v2 client exhibits a ~100% performance drop. This regression affects any user who relies on getXxx() methods and migrates to the new ClickHouse JDBC driver.

I’ve submitted a PR that introduces 2 new benchmarks using strongly-typed getters. Below are the results from running them on my local machine:

  • OpenJDK version: 24.0.2
  • macOS 15.5, Apple M3 Pro
mvn compile exec:exec -Dexec.executable=java -Dexec.args="-classpath %classpath com.clickhouse.benchmark.BenchmarkRunner -m 3 -b q -l 300000"

QueryClient.queryV1                    110  276.252 ± 31.343  ms/op
QueryClient.queryV2                    125  245.277 ± 20.057  ms/op
QueryClient.queryV1WithTypes           144  209.260 ± 17.059  ms/op
QueryClient.queryV2WithTypes           68   454.188 ± 36.329  ms/op

As you can see, QueryClient.queryV1 and QueryClient.queryV2 perform similarly. However, QueryClient.queryV2WithTypes is more than 2x slower than QueryClient.queryV1WithTypes.
While the gap between QueryClient.queryV2 and QueryClient.queryV1WithTypes is around 20%, memory allocations are significantly higher in v2, increasing GC pressure which usually run concurrently.

Root cause

Two main differences in v2 contribute to the regression, both of which are not present in v1:

  • It creates a new Object[] for every row. This means every read triggers an array allocation, and all primitive values are boxed. This increases GC pressure and negatively impacts data locality.
    In contrast, v1 reuses a single array with mutable wrappers to store values, avoiding these allocations entirely (see ClickHouseClientOption#REUSE_VALUE_WRAPPER).
  • In v2, reading a primitive value like getLong(int) involves a chain of unnecessary hash table lookups:
    com.clickhouse.client.api.metadata.TableSchema#columnIndexToName -> nameToIndex -> nameToIndex. com.clickhouse.jdbc.ResultSetImpl#getLong(int) adds one more lookup on top of that chain. This results in 4 HashMap.get() calls per column access by index to read every primitive column value.

Activity

  1. chernser commented on Aug 14, 2025

    @chernser
    Contributor

    Good day, @alekkol !
    Thank you for analysis! We agree with this.

    • Regarding Object[] for each row. This is done because of safety considerations: when value holder is reused then potentially a wrong data can be read from previous row. Some users do not trust this approach. However we will consider having it as an option.
    • Agree with item about index lookup. My mistake - will fix it. So getting fields by index will be most direct addressing to an underlying array.
  2. alekkol commented on Sep 8, 2025

    @alekkol
    ContributorAuthor

    @chernser

    Regarding Object[] for each row. This is done because of safety considerations: when value holder is reused then potentially a wrong data can be read from previous row. Some users do not trust this approach. However we will consider having it as an option.

    Using Object[] forces wrapping of all primitive values. Unfortunately, Java value types are not available yet, so the internal representation is suboptimal. Compressed object headers help a little, but only in recent JVM versions.

    One possible alternative would be a byte-based representation. For example, the PostgreSQL JDBC driver uses byte[][] for rows, where the first index represents a column. This allows primitives to be encoded efficiently (8 bytes for a long, 4 for an int, etc.).

    Just sharing the idea — glad to help!

  3. added this to the 0.9.4 milestone on Oct 1, 2025
  4. removed this from the 0.9.5 milestone on Nov 17, 2025
  5. added this to the 0.9.6 milestone on Dec 18, 2025
  6. modified the milestones: 0.9.6, 0.9.7 on Jan 7, 2026
  7. modified the milestones: 0.9.7, 0.9.8 on Jan 21, 2026
  8. modified the milestones: 0.9.8, 0.9.9 on Mar 16, 2026
  9. dolfinus commented on Mar 31, 2026

    @dolfinus
    Contributor

    On 0.9.8 QueryClient.queryV2 performance is twice as high than QueryClient.queryV1. withTypes is just a bit slower:

    cd ./performance/ &&
    mvn compile exec:exec -Dexec.executable=java -Dexec.args="-classpath %classpath com.clickhouse.benchmark.BenchmarkRunner -m 3 -b q -l 300000"
    Benchmark                                         Cnt          Score        Error   Units
    QueryClient.queryV1                                82        375.333 ±     69.745   ms/op
    QueryClient.queryV2                               155        196.946 ±      9.163   ms/op
    QueryClient.queryV1WithTypes                      122        251.446 ±     32.778   ms/op
    QueryClient.queryV2WithTypes                      101        298.966 ±     35.335   ms/op
    

    GC allocation looks higher, but error is larger than actual measure:

    Benchmark                                         Cnt          Score        Error   Units
    QueryClient.queryV1:gc.alloc.rate                   3        299.173 ±   2029.543  MB/sec
    QueryClient.queryV2:gc.alloc.rate                   3        740.133 ±   1691.498  MB/sec
              
    QueryClient.queryV1:gc.alloc.rate.norm              3  119765750.635 ± 180841.372    B/op
    QueryClient.queryV2:gc.alloc.rate.norm              3  154656992.826 ±  65023.904    B/op
              
    QueryClient.queryV1:gc.count                        3          5.000               counts
    QueryClient.queryV2:gc.count                        3          5.000               counts
              
    QueryClient.queryV1:gc.time                         3         73.000                   ms
    QueryClient.queryV2:gc.time                         3         13.000                   ms
    
    QueryClient.queryV1:mempool.G1 Eden Space.used      3    5025792.000                  KiB
    QueryClient.queryV2:mempool.G1 Eden Space.used      3    5025792.000                  KiB
    
    QueryClient.queryV1:mempool.G1 Old Gen.used         3     323504.000                  KiB
    QueryClient.queryV2:mempool.G1 Old Gen.used         3     327600.000                  KiB
    
    QueryClient.queryV1:mempool.G1 Survivor Space.used  3       3350.984                  KiB
    QueryClient.queryV2:mempool.G1 Survivor Space.used  3       2149.828                  KiB
    
    QueryClient.queryV1:mempool.Metaspace.used          3      25800.609                  KiB
    QueryClient.queryV2:mempool.Metaspace.used          3      25298.477                  KiB
    
    QueryClient.queryV1:mempool.total.used              3    5391359.703                  KiB
    QueryClient.queryV2:mempool.total.used              3    5394522.883                  KiB
    

    Raw data:
    jmh-results-local-1774951135807.json
    jmh-results-local-1774951135807.out.tar.gz

  10. chernser commented on Mar 31, 2026

    @chernser
    Contributor

    @dolfinus

    Thank you for the benchmark!
    I keep this issue open until we implement all improvements.
    Obviously we can do better.

  11. modified the milestones: 0.9.9, 0.9.10 on Apr 15, 2026
  12. Kvel4 commented on Sep 10, 2026

    @Kvel4

    We are seeing what appears to be the same performance problem after migrating the Trino ClickHouse connector from JDBC 0.7.1-patch1 to 0.10.0. For scan-heavy TPC-H queries, execution time increases by approximately 1.94x.

    There is an important difference between this scenario and the QueryClient.queryV2WithTypes benchmark discussed above. That benchmark uses ClickHouseBinaryFormatReader directly, while Trino reads data through JDBC ResultSet.

    The low-level reader supports direct indexed access, but many indexed getters in ResultSetImpl, including getLong(int), getInt(int) and getObject(int), still convert the column index to a name and delegate to the name-based implementation. The name-based path then resolves the name back to an index, sometimes multiple times for the same value.

    Trino has a long-standing JDBC access pattern where it first checks whether a value is NULL:

    resultSet.getObject(columnIndex);
    resultSet.wasNull();

    For a non-null value, Trino then calls the corresponding typed getter:

    resultSet.getLong(columnIndex);

    This pattern was also used with the v1 driver. However, v1 accessed the current row directly by column index, so both operations were relatively cheap. With JDBC v2, the same pattern causes repeated index-to-name and name-to-index conversions, map lookups, and generic object conversion for every non-null value.

    Profiling of the affected workload shows these schema map lookups as a major CPU hotspot, while they are not present with the legacy driver.

    Could you clarify whether completing direct indexed access in JDBC ResultSetImpl is covered by this issue? The remaining indexed getters could use reader.hasValue(columnIndex) and the corresponding reader.getXxx(columnIndex) method directly, similar to the existing implementation of getBytes(int).

  13. chernser commented on Sep 10, 2026

    @chernser
    Contributor

    Good day, @Kvel4 !

    Thank you for bringing this!
    This is planned for upcoming release.

    However I suspect this is not a root cause of slowness.
    Would you please share:

    • what dataset was used for tests (size, structure, number of rows)
    • where is server running, where is benchmark running

    If possible please contribute to performance sub-project in this repository. Thank you!

  14. chernser commented on Sep 10, 2026

    @chernser
    Contributor

    @polyglotAI-bot please implement it. Also update performance project to have JDBC read with and without compression.

  15. Kvel4 commented on Sep 11, 2026

    @Kvel4

    Good day,

    Thank you for the response.

    To be transparent, this was not a fully isolated benchmark environment. The benchmark process, both Trino JVMs, and ClickHouse Server 24.3.14.35 running in Docker were all located on the same host.

    This setup was intentionally chosen as a quick initial check for noticeable performance problems. If needed, I can run a properly isolated benchmark in a cloud environment.

    Both Trino JVMs were running simultaneously, one with ClickHouse JDBC 0.7.1-patch1 and the other with 0.10.0. However, queries were executed against one Trino endpoint at a time, so the second instance was idle during each measurement. Both variants used the same ClickHouse server and the same physical tables.

    Dataset

    The dataset was standard TPC-H at scale factor 4, approximately 4 GB of nominal source data and 34,636,634 rows in total:

    Table Rows
    region 5
    nation 25
    supplier 40,000
    customer 600,000
    part 800,000
    partsupp 3,200,000
    orders 6,000,000
    lineitem 23,996,604

    Benchmark sequence

    Before collecting measurements, the complete suite was executed once as a warm-up for each driver version.

    The measured runs then used 10 alternating blocks with the following sequence in every block:

    0.7.1-patch1 → 0.10.0 → 0.10.0 → 0.7.1-patch1
    

    Each position in this sequence represented one complete execution of the ten-query suite. This produced:

    • 20 measured suite executions per driver version;
    • 20 measurements of every query per driver version;
    • 400 measured query executions in total, excluding warm-ups.

    The benchmark measured end-to-end wall-clock time through Trino.

    Individual SF4 queries took approximately:

    • 0.5–6.9 seconds with JDBC 0.7.1-patch1;
    • 0.8–14.3 seconds with JDBC 0.10.0.

    The median execution time for the complete ten-query suite was approximately:

    • 29.5 seconds with JDBC 0.7.1-patch1;
    • 58.3 seconds with JDBC 0.10.0.

    This corresponds to the approximately 1.94x suite-level degradation mentioned above.

    Host and resource configuration

    • CPU: Intel Core i5-14600KF
    • Architecture: x86_64
    • 20 logical CPUs available
    • RAM: 32 GB
    • Java: OpenJDK 23.0.2
    • Trino JVM heap: -Xms1G -Xmx8G per instance
    • No additional restrictions for docker containers
    • Trino per-node query memory limit: 2.4 GB
  16. chernser commented on Sep 11, 2026

    @chernser
    Contributor

    @Kvel4
    Thank you for sharing the info!

    I will try to reproduce same setup to see how it affects the results. I just have a few thought:

    • running two JVMs + CH server on a single host - how do they interfere with each other.
    • Heap is at 8Gb max and data is 4G - there may be a GC pressure that we need to profile

    What is the complete JVM opts string? is there GC selected?
    Did you monitor with jmxconsole how heap and memory works?
    Do you use JMH? can you share result files with gc allocation and similar information? just for comparison.

    Thank you!

    PS: we started working on the issue.

  17. Kvel4 commented on Sep 11, 2026

    @Kvel4

    Both Trino JVMs and ClickHouse were running on the same host, but only one JVM received queries at a time. The other remained idle. The setup was not fully isolated, but JFR reported approximately 14.9% average host CPU utilization, with a maximum below 47%, so there was no CPU saturation.

    The complete JVM options were:

    -Xms1G -Xmx8G -XX:+UseG1GC -Xlog:gc*,safepoint:file=<file>:time,uptime,level,tags

    Both JVMs used OpenJDK 23.0.2 with G1GC.

    I did not use JConsole. Instead, I recorded both JVMs with JFR using the profile configuration and enabled unified GC logging after the warm-up.

    I also did not use JMH. The benchmark was executed with an internal end-to-end query runner that submits SQL queries to Trino sequentially and saves Trino's server-side query statistics.

    I am attaching the JFR-derived reports and GC logs. The raw JFR recordings are also available if needed.
    jfr-derived-reports.tar.gz

  18. modified the milestones: 0.11.0, 0.12.0 on Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions