Repository navigation
feat(events): persist events to disk so they survive the process (tier 3) - #408
Draft
abelonogov-ld wants to merge 68 commits into
Draft
abelonogov-ld wants to merge 68 commits into
abelonogov-ld wants to merge 68 commits into
Conversation
Tier 3 of the event durability spec: unrecoverable failure. EventStore stages serialized events in memory and commits them to an append-only log in the no-backup files directory, so events survive a process that never gets to flush. Commit points — track, identify, an explicit flush — write on the caller's thread before returning, and the evaluations counted since the last one are summarized and written with them. Delivery reads closed batches back from disk and deletes them only once LaunchDarkly has accepted them, under a payload ID that stays the same across retries so a redelivered batch is not counted twice. Each process of a multi-process application gets its own log. Spec: Event Durability, "Tier 3 — unrecoverable failure"; the mechanism is Page-Cache Event Persistence. Co-authored-by: Cursor <cursoragent@cursor.com>
The counterpart of the iOS EventPersistenceBenchmark, covering the same five context shapes and three privacy settings so the two platforms can be compared rather than assumed alike. Gated on LD_EVENT_BENCH=1, which also turns on stdout and disables up-to-date checks for the test task, so an ordinary run skips all four and is otherwise unaffected. The answer is that Android does not have the problem iOS had. Its unoptimized path beats iOS's fully optimized one on every row: evaluation, summary only 1.87 us (iOS) -> 68 ns evaluation, full event 17.35 us -> 1.51 us evaluation + track 38.40 us -> 4.92 us Cost still tracks attributes written rather than the redaction decision, as on iOS, but at 48-59 ns per attribute against Codable's 1.14 us. So there is no Codable-shaped overhead to remove, because there is no Codable. Co-authored-by: Cursor <cursoragent@cursor.com>
…y/event-durability-tier3-persistence Co-authored-by: Cursor <cursoragent@cursor.com>
…y/event-durability-tier3-persistence
…y/event-durability-tier3-persistence Co-authored-by: Cursor <cursoragent@cursor.com> # Conflicts: # launchdarkly-android-client-sdk/src/main/java/com/launchdarkly/sdk/android/DirectEventProcessor.java
…y/event-durability-tier3-persistence Co-authored-by: Cursor <cursoragent@cursor.com> # Conflicts: # launchdarkly-android-client-sdk/src/main/java/com/launchdarkly/sdk/android/DirectEventProcessor.java # launchdarkly-android-client-sdk/src/main/java/com/launchdarkly/sdk/android/OutboundEventBuffer.java
…y/event-durability-tier3-persistence Co-authored-by: Cursor <cursoragent@cursor.com> # Conflicts: # example/src/main/java/com/launchdarkly/example/MainActivity.java
Coming online now delivers, and the fixture brings the processor online as it builds it, so the queued delivery could drain the store before the test read it. The commit these cover runs offline just the same. Co-authored-by: Cursor <cursoragent@cursor.com>
…y/event-durability-tier3-persistence
…y/event-durability-tier3-persistence Co-authored-by: Cursor <cursoragent@cursor.com> # Conflicts: # launchdarkly-android-client-sdk/src/main/java/com/launchdarkly/sdk/android/ComponentsImpl.java # launchdarkly-android-client-sdk/src/main/java/com/launchdarkly/sdk/android/DirectEventProcessor.java # launchdarkly-android-client-sdk/src/main/java/com/launchdarkly/sdk/android/OutboundEventBuffer.java
Moving these back to where tier 2 has them leaves the capacity fields and their assignment as unchanged lines, so this branch's diff shows the commit machinery it adds rather than a reshuffle of what it inherited. Co-authored-by: Cursor <cursoragent@cursor.com>
…y/event-durability-tier3-persistence Co-authored-by: Cursor <cursoragent@cursor.com> # Conflicts: # example/src/main/java/com/launchdarkly/example/MainActivity.java # example/src/main/res/layout/activity_main.xml
…y/event-durability-tier3-persistence
…y/event-durability-tier3-persistence * andrey/event-durability-tier2-bounded-flush: unserializable unit test
…y/event-durability-tier3-persistence Co-authored-by: Cursor <cursoragent@cursor.com> # Conflicts: # launchdarkly-android-client-sdk/src/main/java/com/launchdarkly/sdk/android/ComponentsImpl.java # launchdarkly-android-client-sdk/src/main/java/com/launchdarkly/sdk/android/DirectEventProcessor.java # launchdarkly-android-client-sdk/src/main/java/com/launchdarkly/sdk/android/OutboundEventBuffer.java
…y/event-durability-tier3-persistence Co-authored-by: Cursor <cursoragent@cursor.com> # Conflicts: # launchdarkly-android-client-sdk/src/main/java/com/launchdarkly/sdk/android/DirectEventProcessor.java
Co-authored-by: Cursor <cursoragent@cursor.com>
…y/event-durability-tier3-persistence * andrey/event-durability-tier2-bounded-flush: test(fdv2): stop requiring a changeset to be the first result
…y/event-durability-tier3-persistence Co-authored-by: Cursor <cursoragent@cursor.com> # Conflicts: # launchdarkly-android-client-sdk/src/main/java/com/launchdarkly/sdk/android/DirectEventProcessor.java # launchdarkly-android-client-sdk/src/test/java/com/launchdarkly/sdk/android/DirectEventProcessorTest.java
At this tier analytics go through AnalyticsEventSender, so the other sender only ever posts diagnostics. Naming it eventSender next to analyticsEventSender read as if it were the general-purpose one. Co-authored-by: Cursor <cursoragent@cursor.com>
…y/event-durability-tier3-persistence Co-authored-by: Cursor <cursoragent@cursor.com>
…y/event-durability-tier3-persistence Co-authored-by: Cursor <cursoragent@cursor.com> # Conflicts: # launchdarkly-android-client-sdk/src/main/java/com/launchdarkly/sdk/android/DirectEventProcessor.java
…y/event-durability-tier3-persistence * andrey/event-durability-tier2-bounded-flush: Document that close() discards events held while offline
…y/event-durability-tier3-persistence Co-authored-by: Cursor <cursoragent@cursor.com> # Conflicts: # launchdarkly-android-client-sdk/src/main/java/com/launchdarkly/sdk/android/DirectEventProcessor.java # launchdarkly-android-client-sdk/src/test/java/com/launchdarkly/sdk/android/EventProcessorTestBase.java
…rent constructor The context limit moved to the fourth parameter; the androidTest copy of the benchmark was never compiled by the unit-test runs and kept the old order. Co-authored-by: Cursor <cursoragent@cursor.com>
A commit takes the pending run under recordLock and encodes it outside, so until it is staged the run is in neither pending nor the store. The capacity check counted only those two, so an event recorded in that window was accepted past capacity; LDClientEventTest.testEventBufferFillsUp saw two identifies with a capacity of one once startup stopped delivering the first one alone. EventStore is no longer final so the test can pause a commit mid-staging. Co-authored-by: Cursor <cursoragent@cursor.com>
With DEFERRED persistence, flush() and blockingFlush() now write on the events thread instead of the calling thread. The event store finds its directory and process name the first time it needs them, on the events thread, not while LDClient.init builds it. close() writes on the events thread and only falls back to the caller's thread if that thread doesn't get to the write within the close budget. commit() returns before taking the I/O lock when persistence is off. The IMMEDIATE and flush() javadocs now say which thread does the write. Co-authored-by: Cursor <cursoragent@cursor.com>
…y/event-durability-tier3-persistence Co-authored-by: Cursor <cursoragent@cursor.com> # Conflicts: # launchdarkly-android-client-sdk/src/main/java/com/launchdarkly/sdk/android/DirectEventProcessor.java
…y/event-durability-tier3-persistence
Tiers 1 and 2 check that a flush never sends an evaluation's counter in one payload and its feature event in the next. Tier 3 dropped that test because a delivery reads the store and can cut or join commits, so payloads no longer show a split. The replacement checks what each commit stages instead: with persistence off every store commit is one of the processor's, and within each the flag's counter must equal its feature events while recording continues. Co-authored-by: Cursor <cursoragent@cursor.com>
With EventPersistence.DISABLED nothing was written, but the store still resolved its directory (which makes Android create no_backup), looked for a previous run's open log at startup, and listed the directory on every delivery. It now skips every read when the application did not ask for persistence. Events a persistent run left on disk stay there until persistence is turned back on. A separate final flag makes the decision, because persistEvents also turns false when a write fails mid-session, and batches written before that failure still have to be read back and delivered. Co-authored-by: Cursor <cursoragent@cursor.com>
Keep tier 3's store-based delivery. Carry tier 1's recording changes into the commit path: RuntimeException guards around recording, the contexts-exceeded warning claimed under recordLock, and the reset of that flag when summaries are taken in commitDurably. With IMMEDIATE persistence the commit runs on the caller's thread, so commitAtCommitPoint and the fallback commit in close() are guarded as well. Co-authored-by: Cursor <cursoragent@cursor.com>
With IMMEDIATE persistence a commit on the caller's thread could open the log before the startup task ran recovery; recovery then closed this run's log as if it were left over, and its events were delivered early as "recovered". Recovery now runs once, under ioLock, before the first append. Co-authored-by: Cursor <cursoragent@cursor.com>
The stream is not in output until the header is written, so a failed header write - a full disk, say - left it open until it was collected. The store gives up on persistence right after, so this leaked at most one descriptor per session. Co-authored-by: Cursor <cursoragent@cursor.com>
describeConfiguration put diagnosticRecordingIntervalMillis twice; the second put only overwrote the first with the same value. Co-authored-by: Cursor <cursoragent@cursor.com>
A commit counted the run it took in eventsBeingStaged until every event was staged, while each staged event was already counted in the store. For the length of the staging loop an event could be counted twice, and one recorded near capacity then was dropped though it would have fit. The run is now reserved in the store in the same critical section it leaves pending in, and staging an event uses up its reservation under the store's lock, so an event is counted in exactly one place throughout. Whatever did not serialize is released when the run is done. Co-authored-by: Cursor <cursoragent@cursor.com>
…w API 24 PerContextEventSummarizer in java-sdk-internal uses Map.computeIfAbsent with a method reference, which D8 turns into a java.util.function.Function; on API 21-23 the first evaluation throws NoClassDefFoundError. Use a copy with a plain get and put until java-sdk-internal ships the same change. Co-authored-by: Cursor <cursoragent@cursor.com>
A scheduled commit (DEFERRED track, the pending threshold) used to queue on the thread that also posts to the network, so a slow post or the one-second sleep before its retry held the commit back for as long. Commits now run on the store's write thread. The retry is scheduled instead of slept for; a flush or close that waits still waits for it within its own budget. Co-authored-by: Cursor <cursoragent@cursor.com>
- Check the context seen last by reference before hashing it. LDContext does not cache its hash, and each evaluation hashed the whole context twice: once for the context limit and once for the per-context summary map. The summarizer shim becomes SameContextPerContextEventSummarizer, which has that fast path and so outlives the API-level fix. - Summarize from the evaluation's values; build a FeatureRequest only for a tracked or debugged flag. - Drop OutboundEventBuffer's monitor, which was only ever taken under DirectEventProcessor's recordLock. - Log an unknown flag at info once per key, and skip building the debug log's arguments when debug logging is off. Co-authored-by: Cursor <cursoragent@cursor.com>
Number every recording, and remember the last one a finished commit covered. A track or identify whose recording a commit already made durable returns without committing again; the next caller still uncovered commits everything that arrived meanwhile. Previously every caller waiting on the commit lock committed in turn, and each of those commits wrote a summary of its own. Co-authored-by: Cursor <cursoragent@cursor.com>
With persistence off a commit writes nothing, as it already skipped for the pending threshold. Now that commits run on their own thread straight away, one queued by identify could land in the middle of the next evaluation's prerequisite walk and split its counters across two summaries. Co-authored-by: Cursor <cursoragent@cursor.com>
Every DEFERRED track asks, from inside its commit point, and the commit thread takes that lock once for each event it stages, so the question queued behind the encode it was only trying to schedule. Co-authored-by: Cursor <cursoragent@cursor.com>
…full run Scheduled commits moved to the store's thread, so waiting behind the events thread only passed when the commit happened to have started. Co-authored-by: Cursor <cursoragent@cursor.com>
…y/event-durability-tier3-persistence
…y/event-durability-tier3-persistence Co-authored-by: Cursor <cursoragent@cursor.com> # Conflicts: # launchdarkly-android-client-sdk/src/main/java/com/launchdarkly/sdk/android/ComponentsImpl.java # launchdarkly-android-client-sdk/src/main/java/com/launchdarkly/sdk/android/DirectEventProcessor.java
…y/event-durability-tier3-persistence Tier 3 keeps a batch the service refuses in a way that may pass, so tier 2's "refused events are lost" does not carry over: the flushAndWait javadoc says they are retried, and the tier-2 tests that pinned the loss now pin that the batch is kept and resent under its payload ID once the service recovers. The lost-events flag answers only for real losses here: an event or summary that could not be serialized at commit, reported by OutboundEventBuffer. Adds tests that a refused batch survives repeated flushes past the standard retry, for every recoverable status, in memory and on disk, alongside newer batches, and across a restart. Co-authored-by: Cursor <cursoragent@cursor.com>
…retry Co-authored-by: Cursor <cursoragent@cursor.com>
…y/event-durability-tier3-persistence
Co-authored-by: Cursor <cursoragent@cursor.com>
…ents too large to store closeBatch closed only the log when a failed write left events both there and in memory, so a flush could succeed and close() discard the rest. A frame over the size limit was dropped without telling the flush. Startup recovery could miss an online transition and wait for the periodic flush. Co-authored-by: Cursor <cursoragent@cursor.com>
A read can fail on an intact file, as it does when the process is out of file descriptors, and every failure was treated as an unreadable batch and deleted. Listing now skips such a batch until a later listing can read it, delivery retries it, and startup recovery moves the previous run's log aside uncounted rather than deleting it. Only a missing file reads as nothing to send. Co-authored-by: Cursor <cursoragent@cursor.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Requirements
Related issues
Tier 3 of the Event Durability spec, "Tier 3 — persistence". Stacked on #403 (tier 2) and targets that branch, so this diff shows only tier 3. It will be retargeted once the branches below it merge. The iOS counterpart is launchdarkly/ios-client-sdk#538, and both write the same on-disk format.
Opened as a draft: a multi-threading and readability review turned up items still being worked through (see Known open items).
Describe the solution you've provided
Recorded events go to an append-only log on disk rather than only an in-memory buffer, so an application that dies shortly after recording an event still reports it: the next run delivers what the previous one left behind. That covers what tier 2 cannot:
SIGKILL, ANR kills, native crashes, and the system reclaiming a backgrounded process.EventPersistence(new public enum) andEventProcessorBuilder.eventPersistence(...):DISABLED(default): events stay in memory and the SDK does not touch the filesystem for events.DEFERRED: writes happen on the SDK's commit thread, a moment after recording.IMMEDIATE:trackandidentifywrite on the caller's thread before returning, so there is no window.EventStore: an append-only log under the no-backup files directory, one directory per mobile key. It handles framing, capacity, torn-tail recovery of a log interrupted by a crash, and a fallback to memory when a write fails (for example, a full disk), so a disk failure never takes the application down. Each process appends to its ownopen-<process>log; closedready-<payloadId>batches are shared, so a batch left by a process that never runs again is still delivered.DirectEventProcessor): an evaluation appends to a pending run and updates a counter, with no encoding on the caller's thread. A commit encodes the run and its summaries together and stages them in the store.trackandidentifyare commit points. Evaluations commit once a threshold accumulates, on flush, and before delivery.AnalyticsEventSenderposts a pre-assembled body under an explicit payload ID.DEFAULT_CAPACITYgoes from 100 to 1000, matching iOS, since capacity now also bounds what is held on disk.Describe alternatives you've considered
SQLite. Measured as a store and rejected: it costs more per event than an append to a log the SDK already frames, for durability properties the log gets from
writealone. Seesqlite-event-persistence-findings.mdin the spec folder.Fsync on every commit. What durability is defending against is the process dying, not the device losing power. Bytes handed to the kernel survive the former, and an fsync per event would dominate the cost of
track.One shared log for all processes. That is simpler, but one process's startup recovery could rename a log another live process is still appending to. A log per process avoids that without file locks.
Additional context
Tests. 873 unit tests pass locally (
testDebugUnitTest). New coverage includesEventStoreMultiProcessTest(several stores on one directory, recovery, transient read failures),DirectEventProcessorTest(commit points, retry, capacity across memory and disk, persistence failing partway through, oversized frames reported as losses),OutboundEventBufferSerializationTest, andSameContextPerContextEventSummarizerTest.Benchmarks.
EventPersistenceBenchmark(JVM and instrumented) measures recording cost per mode. It runs only whenLD_EVENT_BENCH=1, so ordinary test runs are unaffected.Test app. Now configures
EventPersistence.IMMEDIATE, so the events recorded by Eval+Kill now, which tier 2 loses, are recovered and delivered on the next launch. It also gains Over-Refresh Eval, a main-thread stress loop that compares evaluation cost while filling capacity and after it is full.Known open items (from the threading review, to be resolved before this leaves draft):
flushAsync()can be consumed by a different flush than the one that lost them.close()'s fallback commit can wait on the commit lock past the two-second close budget if the disk stalls.OutboundEventBufferare now unused outside tests, and a few comments still describe tier 2.Version.
EventPersistenceandEventProcessorBuilder.eventPersistencehave no@sincetag yet; one should be added once the release version is known.