Skip to content

feat(events): persist events to disk so they survive the process (tier 3) - #408

Draft
abelonogov-ld wants to merge 68 commits into
andrey/event-durability-tier2-bounded-flushfrom
andrey/event-durability-tier3-persistence
Draft

abelonogov-ld wants to merge 68 commits into
andrey/event-durability-tier2-bounded-flushfrom
andrey/event-durability-tier3-persistence

Conversation

@abelonogov-ld

Copy link
Copy Markdown
Contributor

Requirements

  • I have added test coverage for new or changed functionality
  • I have followed the repository's pull request submission guidelines
  • I have validated my changes against all supported platform versions

Related issues

Tier 3 of the Event Durability spec, "Tier 3 — persistence". Stacked on #403 (tier 2) and targets that branch, so this diff shows only tier 3. It will be retargeted once the branches below it merge. The iOS counterpart is launchdarkly/ios-client-sdk#538, and both write the same on-disk format.

Opened as a draft: a multi-threading and readability review turned up items still being worked through (see Known open items).

Describe the solution you've provided

Recorded events go to an append-only log on disk rather than only an in-memory buffer, so an application that dies shortly after recording an event still reports it: the next run delivers what the previous one left behind. That covers what tier 2 cannot: SIGKILL, ANR kills, native crashes, and the system reclaiming a backgrounded process.

  • EventPersistence (new public enum) and EventProcessorBuilder.eventPersistence(...):
    • DISABLED (default): events stay in memory and the SDK does not touch the filesystem for events.
    • DEFERRED: writes happen on the SDK's commit thread, a moment after recording.
    • IMMEDIATE: track and identify write on the caller's thread before returning, so there is no window.
  • EventStore: an append-only log under the no-backup files directory, one directory per mobile key. It handles framing, capacity, torn-tail recovery of a log interrupted by a crash, and a fallback to memory when a write fails (for example, a full disk), so a disk failure never takes the application down. Each process appends to its own open-<process> log; closed ready-<payloadId> batches are shared, so a batch left by a process that never runs again is still delivered.
  • Recording and commits (DirectEventProcessor): an evaluation appends to a pending run and updates a counter, with no encoding on the caller's thread. A commit encodes the run and its summaries together and stages them in the store. track and identify are commit points. Evaluations commit once a threshold accumulates, on flush, and before delivery.
  • Delivery reads batches oldest first. A batch keeps its payload ID across retries and process restarts, so LaunchDarkly discards a repeated delivery instead of double counting. Delivery stops at the first retryable failure, schedules a single retry instead of sleeping on the events thread, and removes a batch only once it has been accepted or permanently refused. A batch whose read fails transiently is kept for a later attempt.
  • AnalyticsEventSender posts a pre-assembled body under an explicit payload ID.
  • DEFAULT_CAPACITY goes from 100 to 1000, matching iOS, since capacity now also bounds what is held on disk.

Describe alternatives you've considered

SQLite. Measured as a store and rejected: it costs more per event than an append to a log the SDK already frames, for durability properties the log gets from write alone. See sqlite-event-persistence-findings.md in the spec folder.

Fsync on every commit. What durability is defending against is the process dying, not the device losing power. Bytes handed to the kernel survive the former, and an fsync per event would dominate the cost of track.

One shared log for all processes. That is simpler, but one process's startup recovery could rename a log another live process is still appending to. A log per process avoids that without file locks.

Additional context

Tests. 873 unit tests pass locally (testDebugUnitTest). New coverage includes EventStoreMultiProcessTest (several stores on one directory, recovery, transient read failures), DirectEventProcessorTest (commit points, retry, capacity across memory and disk, persistence failing partway through, oversized frames reported as losses), OutboundEventBufferSerializationTest, and SameContextPerContextEventSummarizerTest.

Benchmarks. EventPersistenceBenchmark (JVM and instrumented) measures recording cost per mode. It runs only when LD_EVENT_BENCH=1, so ordinary test runs are unaffected.

Test app. Now configures EventPersistence.IMMEDIATE, so the events recorded by Eval+Kill now, which tier 2 loses, are recovered and delivered on the next launch. It also gains Over-Refresh Eval, a main-thread stress loop that compares evaluation cost while filling capacity and after it is full.

Known open items (from the threading review, to be resolved before this leaves draft):

  • Capacity briefly under-counts a batch while it is being written to disk.
  • The "events were lost" flag behind flushAsync() can be consumed by a different flush than the one that lost them.
  • close()'s fallback commit can wait on the commit lock past the two-second close budget if the disk stalls.
  • Batches left by a previous run are not counted against capacity until the first listing, which an offline start postpones.
  • Readability: about 200 lines of tier-2 whole-payload encoding in OutboundEventBuffer are now unused outside tests, and a few comments still describe tier 2.

Version. EventPersistence and EventProcessorBuilder.eventPersistence have no @since tag yet; one should be added once the release version is known.

abelonogov-ld and others added 30 commits September 9, 2026 11:59
Tier 3 of the event durability spec: unrecoverable failure.

EventStore stages serialized events in memory and commits them to an append-only
log in the no-backup files directory, so events survive a process that never gets
to flush. Commit points — track, identify, an explicit flush — write on the
caller's thread before returning, and the evaluations counted since the last one
are summarized and written with them. Delivery reads closed batches back from
disk and deletes them only once LaunchDarkly has accepted them, under a payload
ID that stays the same across retries so a redelivered batch is not counted
twice. Each process of a multi-process application gets its own log.

Spec: Event Durability, "Tier 3 — unrecoverable failure"; the mechanism is
Page-Cache Event Persistence.

Co-authored-by: Cursor <cursoragent@cursor.com>
The counterpart of the iOS EventPersistenceBenchmark, covering the same
five context shapes and three privacy settings so the two platforms can
be compared rather than assumed alike. Gated on LD_EVENT_BENCH=1, which
also turns on stdout and disables up-to-date checks for the test task,
so an ordinary run skips all four and is otherwise unaffected.

The answer is that Android does not have the problem iOS had. Its
unoptimized path beats iOS's fully optimized one on every row:

  evaluation, summary only      1.87 us (iOS) ->  68 ns
  evaluation, full event       17.35 us       -> 1.51 us
  evaluation + track            38.40 us       -> 4.92 us

Cost still tracks attributes written rather than the redaction decision,
as on iOS, but at 48-59 ns per attribute against Codable's 1.14 us. So
there is no Codable-shaped overhead to remove, because there is no
Codable.

Co-authored-by: Cursor <cursoragent@cursor.com>
…y/event-durability-tier3-persistence

Co-authored-by: Cursor <cursoragent@cursor.com>
…y/event-durability-tier3-persistence

Co-authored-by: Cursor <cursoragent@cursor.com>

# Conflicts:
#	launchdarkly-android-client-sdk/src/main/java/com/launchdarkly/sdk/android/DirectEventProcessor.java
…y/event-durability-tier3-persistence

Co-authored-by: Cursor <cursoragent@cursor.com>

# Conflicts:
#	launchdarkly-android-client-sdk/src/main/java/com/launchdarkly/sdk/android/DirectEventProcessor.java
#	launchdarkly-android-client-sdk/src/main/java/com/launchdarkly/sdk/android/OutboundEventBuffer.java
…y/event-durability-tier3-persistence

Co-authored-by: Cursor <cursoragent@cursor.com>

# Conflicts:
#	example/src/main/java/com/launchdarkly/example/MainActivity.java
Coming online now delivers, and the fixture brings the processor online as
it builds it, so the queued delivery could drain the store before the test
read it. The commit these cover runs offline just the same.

Co-authored-by: Cursor <cursoragent@cursor.com>
…y/event-durability-tier3-persistence

Co-authored-by: Cursor <cursoragent@cursor.com>

# Conflicts:
#	launchdarkly-android-client-sdk/src/main/java/com/launchdarkly/sdk/android/ComponentsImpl.java
#	launchdarkly-android-client-sdk/src/main/java/com/launchdarkly/sdk/android/DirectEventProcessor.java
#	launchdarkly-android-client-sdk/src/main/java/com/launchdarkly/sdk/android/OutboundEventBuffer.java
Moving these back to where tier 2 has them leaves the capacity fields and
their assignment as unchanged lines, so this branch's diff shows the commit
machinery it adds rather than a reshuffle of what it inherited.

Co-authored-by: Cursor <cursoragent@cursor.com>
…y/event-durability-tier3-persistence

Co-authored-by: Cursor <cursoragent@cursor.com>

# Conflicts:
#	example/src/main/java/com/launchdarkly/example/MainActivity.java
#	example/src/main/res/layout/activity_main.xml
…y/event-durability-tier3-persistence

* andrey/event-durability-tier2-bounded-flush:
  unserializable unit test
…y/event-durability-tier3-persistence

Co-authored-by: Cursor <cursoragent@cursor.com>

# Conflicts:
#	launchdarkly-android-client-sdk/src/main/java/com/launchdarkly/sdk/android/ComponentsImpl.java
#	launchdarkly-android-client-sdk/src/main/java/com/launchdarkly/sdk/android/DirectEventProcessor.java
#	launchdarkly-android-client-sdk/src/main/java/com/launchdarkly/sdk/android/OutboundEventBuffer.java
…y/event-durability-tier3-persistence

Co-authored-by: Cursor <cursoragent@cursor.com>

# Conflicts:
#	launchdarkly-android-client-sdk/src/main/java/com/launchdarkly/sdk/android/DirectEventProcessor.java
Co-authored-by: Cursor <cursoragent@cursor.com>
…y/event-durability-tier3-persistence

* andrey/event-durability-tier2-bounded-flush:
  test(fdv2): stop requiring a changeset to be the first result
…y/event-durability-tier3-persistence

Co-authored-by: Cursor <cursoragent@cursor.com>

# Conflicts:
#	launchdarkly-android-client-sdk/src/main/java/com/launchdarkly/sdk/android/DirectEventProcessor.java
#	launchdarkly-android-client-sdk/src/test/java/com/launchdarkly/sdk/android/DirectEventProcessorTest.java
At this tier analytics go through AnalyticsEventSender, so the other
sender only ever posts diagnostics. Naming it eventSender next to
analyticsEventSender read as if it were the general-purpose one.

Co-authored-by: Cursor <cursoragent@cursor.com>
…y/event-durability-tier3-persistence

Co-authored-by: Cursor <cursoragent@cursor.com>
…y/event-durability-tier3-persistence

Co-authored-by: Cursor <cursoragent@cursor.com>

# Conflicts:
#	launchdarkly-android-client-sdk/src/main/java/com/launchdarkly/sdk/android/DirectEventProcessor.java
abelonogov-ld and others added 30 commits September 22, 2026 18:37
…y/event-durability-tier3-persistence

* andrey/event-durability-tier2-bounded-flush:
  Document that close() discards events held while offline
…y/event-durability-tier3-persistence

Co-authored-by: Cursor <cursoragent@cursor.com>

# Conflicts:
#	launchdarkly-android-client-sdk/src/main/java/com/launchdarkly/sdk/android/DirectEventProcessor.java
#	launchdarkly-android-client-sdk/src/test/java/com/launchdarkly/sdk/android/EventProcessorTestBase.java
…rent constructor

The context limit moved to the fourth parameter; the androidTest copy of the benchmark was never
compiled by the unit-test runs and kept the old order.

Co-authored-by: Cursor <cursoragent@cursor.com>
A commit takes the pending run under recordLock and encodes it outside, so until it is staged the run
is in neither pending nor the store. The capacity check counted only those two, so an event recorded in
that window was accepted past capacity; LDClientEventTest.testEventBufferFillsUp saw two identifies
with a capacity of one once startup stopped delivering the first one alone. EventStore is no longer
final so the test can pause a commit mid-staging.

Co-authored-by: Cursor <cursoragent@cursor.com>
With DEFERRED persistence, flush() and blockingFlush() now write on the
events thread instead of the calling thread. The event store finds its
directory and process name the first time it needs them, on the events
thread, not while LDClient.init builds it. close() writes on the events
thread and only falls back to the caller's thread if that thread doesn't
get to the write within the close budget. commit() returns before taking
the I/O lock when persistence is off. The IMMEDIATE and flush() javadocs
now say which thread does the write.

Co-authored-by: Cursor <cursoragent@cursor.com>
…y/event-durability-tier3-persistence

Co-authored-by: Cursor <cursoragent@cursor.com>

# Conflicts:
#	launchdarkly-android-client-sdk/src/main/java/com/launchdarkly/sdk/android/DirectEventProcessor.java
Tiers 1 and 2 check that a flush never sends an evaluation's counter in
one payload and its feature event in the next. Tier 3 dropped that test
because a delivery reads the store and can cut or join commits, so
payloads no longer show a split. The replacement checks what each
commit stages instead: with persistence off every store commit is one
of the processor's, and within each the flag's counter must equal its
feature events while recording continues.

Co-authored-by: Cursor <cursoragent@cursor.com>
With EventPersistence.DISABLED nothing was written, but the store still
resolved its directory (which makes Android create no_backup), looked
for a previous run's open log at startup, and listed the directory on
every delivery. It now skips every read when the application did not
ask for persistence. Events a persistent run left on disk stay there
until persistence is turned back on.

A separate final flag makes the decision, because persistEvents also
turns false when a write fails mid-session, and batches written before
that failure still have to be read back and delivered.

Co-authored-by: Cursor <cursoragent@cursor.com>
Keep tier 3's store-based delivery. Carry tier 1's recording changes into
the commit path: RuntimeException guards around recording, the
contexts-exceeded warning claimed under recordLock, and the reset of that
flag when summaries are taken in commitDurably. With IMMEDIATE persistence
the commit runs on the caller's thread, so commitAtCommitPoint and the
fallback commit in close() are guarded as well.

Co-authored-by: Cursor <cursoragent@cursor.com>
With IMMEDIATE persistence a commit on the caller's thread could open the
log before the startup task ran recovery; recovery then closed this run's
log as if it were left over, and its events were delivered early as
"recovered". Recovery now runs once, under ioLock, before the first append.

Co-authored-by: Cursor <cursoragent@cursor.com>
The stream is not in output until the header is written, so a failed
header write - a full disk, say - left it open until it was collected.
The store gives up on persistence right after, so this leaked at most one
descriptor per session.

Co-authored-by: Cursor <cursoragent@cursor.com>
describeConfiguration put diagnosticRecordingIntervalMillis twice; the
second put only overwrote the first with the same value.

Co-authored-by: Cursor <cursoragent@cursor.com>
A commit counted the run it took in eventsBeingStaged until every event
was staged, while each staged event was already counted in the store. For
the length of the staging loop an event could be counted twice, and one
recorded near capacity then was dropped though it would have fit.

The run is now reserved in the store in the same critical section it
leaves pending in, and staging an event uses up its reservation under the
store's lock, so an event is counted in exactly one place throughout.
Whatever did not serialize is released when the run is done.

Co-authored-by: Cursor <cursoragent@cursor.com>
…w API 24

PerContextEventSummarizer in java-sdk-internal uses Map.computeIfAbsent with a
method reference, which D8 turns into a java.util.function.Function; on API 21-23
the first evaluation throws NoClassDefFoundError. Use a copy with a plain get and
put until java-sdk-internal ships the same change.

Co-authored-by: Cursor <cursoragent@cursor.com>
A scheduled commit (DEFERRED track, the pending threshold) used to queue on the
thread that also posts to the network, so a slow post or the one-second sleep
before its retry held the commit back for as long. Commits now run on the
store's write thread. The retry is scheduled instead of slept for; a flush or
close that waits still waits for it within its own budget.

Co-authored-by: Cursor <cursoragent@cursor.com>
- Check the context seen last by reference before hashing it. LDContext does
  not cache its hash, and each evaluation hashed the whole context twice: once
  for the context limit and once for the per-context summary map. The
  summarizer shim becomes SameContextPerContextEventSummarizer, which has that
  fast path and so outlives the API-level fix.
- Summarize from the evaluation's values; build a FeatureRequest only for a
  tracked or debugged flag.
- Drop OutboundEventBuffer's monitor, which was only ever taken under
  DirectEventProcessor's recordLock.
- Log an unknown flag at info once per key, and skip building the debug log's
  arguments when debug logging is off.

Co-authored-by: Cursor <cursoragent@cursor.com>
Number every recording, and remember the last one a finished commit covered. A
track or identify whose recording a commit already made durable returns without
committing again; the next caller still uncovered commits everything that
arrived meanwhile. Previously every caller waiting on the commit lock committed
in turn, and each of those commits wrote a summary of its own.

Co-authored-by: Cursor <cursoragent@cursor.com>
With persistence off a commit writes nothing, as it already skipped for the
pending threshold. Now that commits run on their own thread straight away,
one queued by identify could land in the middle of the next evaluation's
prerequisite walk and split its counters across two summaries.

Co-authored-by: Cursor <cursoragent@cursor.com>
Every DEFERRED track asks, from inside its commit point, and the commit
thread takes that lock once for each event it stages, so the question
queued behind the encode it was only trying to schedule.

Co-authored-by: Cursor <cursoragent@cursor.com>
…full run

Scheduled commits moved to the store's thread, so waiting behind the
events thread only passed when the commit happened to have started.

Co-authored-by: Cursor <cursoragent@cursor.com>
…y/event-durability-tier3-persistence

Co-authored-by: Cursor <cursoragent@cursor.com>

# Conflicts:
#	launchdarkly-android-client-sdk/src/main/java/com/launchdarkly/sdk/android/ComponentsImpl.java
#	launchdarkly-android-client-sdk/src/main/java/com/launchdarkly/sdk/android/DirectEventProcessor.java
…y/event-durability-tier3-persistence

Tier 3 keeps a batch the service refuses in a way that may pass, so tier 2's
"refused events are lost" does not carry over: the flushAndWait javadoc says
they are retried, and the tier-2 tests that pinned the loss now pin that the
batch is kept and resent under its payload ID once the service recovers.

The lost-events flag answers only for real losses here: an event or summary
that could not be serialized at commit, reported by OutboundEventBuffer.

Adds tests that a refused batch survives repeated flushes past the standard
retry, for every recoverable status, in memory and on disk, alongside newer
batches, and across a restart.

Co-authored-by: Cursor <cursoragent@cursor.com>
…retry

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
…ents too large to store

closeBatch closed only the log when a failed write left events both there and in memory, so a
flush could succeed and close() discard the rest. A frame over the size limit was dropped without
telling the flush. Startup recovery could miss an online transition and wait for the periodic flush.

Co-authored-by: Cursor <cursoragent@cursor.com>
A read can fail on an intact file, as it does when the process is out of file descriptors, and
every failure was treated as an unreadable batch and deleted. Listing now skips such a batch until
a later listing can read it, delivery retries it, and startup recovery moves the previous run's log
aside uncounted rather than deleting it. Only a missing file reads as nothing to send.

Co-authored-by: Cursor <cursoragent@cursor.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant