Repository navigation
Major performance regression in client-v2 #2516
Description
Activity
Good day, @alekkol !
Thank you for analysis! We agree with this.- Regarding Object[] for each row. This is done because of safety considerations: when value holder is reused then potentially a wrong data can be read from previous row. Some users do not trust this approach. However we will consider having it as an option.
- Agree with item about index lookup. My mistake - will fix it. So getting fields by index will be most direct addressing to an underlying array.
Regarding Object[] for each row. This is done because of safety considerations: when value holder is reused then potentially a wrong data can be read from previous row. Some users do not trust this approach. However we will consider having it as an option.
Using
Object[]forces wrapping of all primitive values. Unfortunately, Java value types are not available yet, so the internal representation is suboptimal. Compressed object headers help a little, but only in recent JVM versions.One possible alternative would be a byte-based representation. For example, the PostgreSQL JDBC driver uses
byte[][]for rows, where the first index represents a column. This allows primitives to be encoded efficiently (8 bytes for a long, 4 for an int, etc.).Just sharing the idea — glad to help!
Reacted by Sergey ChernovOn 0.9.8
QueryClient.queryV2performance is twice as high thanQueryClient.queryV1.withTypesis just a bit slower:cd ./performance/ && mvn compile exec:exec -Dexec.executable=java -Dexec.args="-classpath %classpath com.clickhouse.benchmark.BenchmarkRunner -m 3 -b q -l 300000"
Benchmark Cnt Score Error Units QueryClient.queryV1 82 375.333 ± 69.745 ms/op QueryClient.queryV2 155 196.946 ± 9.163 ms/op QueryClient.queryV1WithTypes 122 251.446 ± 32.778 ms/op QueryClient.queryV2WithTypes 101 298.966 ± 35.335 ms/opGC allocation looks higher, but error is larger than actual measure:
Benchmark Cnt Score Error Units QueryClient.queryV1:gc.alloc.rate 3 299.173 ± 2029.543 MB/sec QueryClient.queryV2:gc.alloc.rate 3 740.133 ± 1691.498 MB/sec QueryClient.queryV1:gc.alloc.rate.norm 3 119765750.635 ± 180841.372 B/op QueryClient.queryV2:gc.alloc.rate.norm 3 154656992.826 ± 65023.904 B/op QueryClient.queryV1:gc.count 3 5.000 counts QueryClient.queryV2:gc.count 3 5.000 counts QueryClient.queryV1:gc.time 3 73.000 ms QueryClient.queryV2:gc.time 3 13.000 ms QueryClient.queryV1:mempool.G1 Eden Space.used 3 5025792.000 KiB QueryClient.queryV2:mempool.G1 Eden Space.used 3 5025792.000 KiB QueryClient.queryV1:mempool.G1 Old Gen.used 3 323504.000 KiB QueryClient.queryV2:mempool.G1 Old Gen.used 3 327600.000 KiB QueryClient.queryV1:mempool.G1 Survivor Space.used 3 3350.984 KiB QueryClient.queryV2:mempool.G1 Survivor Space.used 3 2149.828 KiB QueryClient.queryV1:mempool.Metaspace.used 3 25800.609 KiB QueryClient.queryV2:mempool.Metaspace.used 3 25298.477 KiB QueryClient.queryV1:mempool.total.used 3 5391359.703 KiB QueryClient.queryV2:mempool.total.used 3 5394522.883 KiBRaw data:
jmh-results-local-1774951135807.json
jmh-results-local-1774951135807.out.tar.gzThank you for the benchmark!
I keep this issue open until we implement all improvements.
Obviously we can do better.Reacted by Maxim Martynov and Alexander KolesnikovWe are seeing what appears to be the same performance problem after migrating the Trino ClickHouse connector from JDBC
0.7.1-patch1to0.10.0. For scan-heavy TPC-H queries, execution time increases by approximately1.94x.There is an important difference between this scenario and the
QueryClient.queryV2WithTypesbenchmark discussed above. That benchmark usesClickHouseBinaryFormatReaderdirectly, while Trino reads data through JDBCResultSet.The low-level reader supports direct indexed access, but many indexed getters in
ResultSetImpl, includinggetLong(int),getInt(int)andgetObject(int), still convert the column index to a name and delegate to the name-based implementation. The name-based path then resolves the name back to an index, sometimes multiple times for the same value.Trino has a long-standing JDBC access pattern where it first checks whether a value is NULL:
resultSet.getObject(columnIndex); resultSet.wasNull();
For a non-null value, Trino then calls the corresponding typed getter:
resultSet.getLong(columnIndex);
This pattern was also used with the v1 driver. However, v1 accessed the current row directly by column index, so both operations were relatively cheap. With JDBC v2, the same pattern causes repeated index-to-name and name-to-index conversions, map lookups, and generic object conversion for every non-null value.
Profiling of the affected workload shows these schema map lookups as a major CPU hotspot, while they are not present with the legacy driver.
Could you clarify whether completing direct indexed access in JDBC
ResultSetImplis covered by this issue? The remaining indexed getters could usereader.hasValue(columnIndex)and the correspondingreader.getXxx(columnIndex)method directly, similar to the existing implementation ofgetBytes(int).Reacted by Sergey ChernovGood day, @Kvel4 !
Thank you for bringing this!
This is planned for upcoming release.However I suspect this is not a root cause of slowness.
Would you please share:- what dataset was used for tests (size, structure, number of rows)
- where is server running, where is benchmark running
If possible please contribute to
performancesub-project in this repository. Thank you!@polyglotAI-bot please implement it. Also update
performanceproject to have JDBC read with and without compression.Good day,
Thank you for the response.
To be transparent, this was not a fully isolated benchmark environment. The benchmark process, both Trino JVMs, and ClickHouse Server
24.3.14.35running in Docker were all located on the same host.This setup was intentionally chosen as a quick initial check for noticeable performance problems. If needed, I can run a properly isolated benchmark in a cloud environment.
Both Trino JVMs were running simultaneously, one with ClickHouse JDBC
0.7.1-patch1and the other with0.10.0. However, queries were executed against one Trino endpoint at a time, so the second instance was idle during each measurement. Both variants used the same ClickHouse server and the same physical tables.Dataset
The dataset was standard TPC-H at scale factor 4, approximately 4 GB of nominal source data and 34,636,634 rows in total:
Table Rows region5 nation25 supplier40,000 customer600,000 part800,000 partsupp3,200,000 orders6,000,000 lineitem23,996,604 Benchmark sequence
Before collecting measurements, the complete suite was executed once as a warm-up for each driver version.
The measured runs then used 10 alternating blocks with the following sequence in every block:
0.7.1-patch1 → 0.10.0 → 0.10.0 → 0.7.1-patch1Each position in this sequence represented one complete execution of the ten-query suite. This produced:
- 20 measured suite executions per driver version;
- 20 measurements of every query per driver version;
- 400 measured query executions in total, excluding warm-ups.
The benchmark measured end-to-end wall-clock time through Trino.
Individual SF4 queries took approximately:
0.5–6.9 secondswith JDBC0.7.1-patch1;0.8–14.3 secondswith JDBC0.10.0.
The median execution time for the complete ten-query suite was approximately:
29.5 secondswith JDBC0.7.1-patch1;58.3 secondswith JDBC0.10.0.
This corresponds to the approximately
1.94xsuite-level degradation mentioned above.Host and resource configuration
- CPU: Intel Core i5-14600KF
- Architecture: x86_64
- 20 logical CPUs available
- RAM: 32 GB
- Java: OpenJDK 23.0.2
- Trino JVM heap:
-Xms1G -Xmx8Gper instance - No additional restrictions for docker containers
- Trino per-node query memory limit: 2.4 GB
@Kvel4
Thank you for sharing the info!I will try to reproduce same setup to see how it affects the results. I just have a few thought:
- running two JVMs + CH server on a single host - how do they interfere with each other.
- Heap is at 8Gb max and data is 4G - there may be a GC pressure that we need to profile
What is the complete JVM opts string? is there GC selected?
Did you monitor with jmxconsole how heap and memory works?
Do you use JMH? can you share result files with gc allocation and similar information? just for comparison.Thank you!
PS: we started working on the issue.
Both Trino JVMs and ClickHouse were running on the same host, but only one JVM received queries at a time. The other remained idle. The setup was not fully isolated, but JFR reported approximately 14.9% average host CPU utilization, with a maximum below 47%, so there was no CPU saturation.
The complete JVM options were:
-Xms1G -Xmx8G -XX:+UseG1GC -Xlog:gc*,safepoint:file=<file>:time,uptime,level,tagsBoth JVMs used OpenJDK 23.0.2 with G1GC.
I did not use JConsole. Instead, I recorded both JVMs with JFR using the
profileconfiguration and enabled unified GC logging after the warm-up.I also did not use JMH. The benchmark was executed with an internal end-to-end query runner that submits SQL queries to Trino sequentially and saves Trino's server-side query statistics.
I am attaching the JFR-derived reports and GC logs. The raw JFR recordings are also available if needed.
jfr-derived-reports.tar.gz
Summary
After migrating from the ClickHouse Java client v1 to v2 we observed a major performance regression: more than 2 times less throughput. We use ClickHouse as a pre-aggregation layer, and run analytical queries that may return 10^6-10^9 rows for later processing. Because our pipeline depends on low-latency, low-overhead reads, even small regressions translate into major throughput losses.
Reproduction
The benchmark suite in this repository measures the general query performance differences between v1 and v2 but does not cover the most performance-sensitive scenario: retrieving column values using strongly-typed getters (e.g., the reason
java.sql.ResultSet#getLong()exists). In such cases, the new v2 client exhibits a ~100% performance drop. This regression affects any user who relies ongetXxx()methods and migrates to the new ClickHouse JDBC driver.I’ve submitted a PR that introduces 2 new benchmarks using strongly-typed getters. Below are the results from running them on my local machine:
As you can see,
QueryClient.queryV1andQueryClient.queryV2perform similarly. However,QueryClient.queryV2WithTypesis more than 2x slower thanQueryClient.queryV1WithTypes.While the gap between
QueryClient.queryV2andQueryClient.queryV1WithTypesis around 20%, memory allocations are significantly higher in v2, increasing GC pressure which usually run concurrently.Root cause
Two main differences in v2 contribute to the regression, both of which are not present in v1:
Object[]for every row. This means every read triggers an array allocation, and all primitive values are boxed. This increases GC pressure and negatively impacts data locality.In contrast, v1 reuses a single array with mutable wrappers to store values, avoiding these allocations entirely (see
ClickHouseClientOption#REUSE_VALUE_WRAPPER).getLong(int)involves a chain of unnecessary hash table lookups:com.clickhouse.client.api.metadata.TableSchema#columnIndexToName->nameToIndex->nameToIndex.com.clickhouse.jdbc.ResultSetImpl#getLong(int)adds one more lookup on top of that chain. This results in 4HashMap.get()calls per column access by index to read every primitive column value.