Files
nfsmw-online/docs/BETA_TELEMETRY_PLAN.md
T
megboyzzandClaude e7c76fc2dd docs: bring 1.8 MB of project documentation under version control
These files had never been tracked anywhere - they lived in a plain directory
with no git at all, which is also where the whole reverse-engineering record
sat. Code already committed refers to them by name (opponent_substitution.h
cites "ANALYSIS.md section 6hh", DebugMenuOverlay.kt cites "DEBUG_MENU.md
section 3"), so until now a fresh clone carried references to documents it did
not contain.

  ANALYSIS.md                       the RE record, and the reason the rest works
  ARCHITECTURE.md                   how the mod's pieces fit together
  ARM64_TRANSLATION_LAYER.md        the translation layer's running log
  PROGRESS.md                       chronological progress across both chats
  BETA_TELEMETRY_PLAN.md            how crash/telemetry reporting is meant to work
  LOBBY_UI_DESIGN.md + .html        lobby design and its clickable prototype
  DEBUG_MENU.md                     debug panel design
  STATIC_RECOMPILATION_FALLBACK.md  the plan if translation had not panned out
  evidence/                         font atlas capture from the glyph-corruption bug
  save_backups/                     saves at known milestones, for reproducing state

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-22 23:29:12 +03:00

14 KiB

Closed beta: diagnostics, crash reports and tester workflow

Plan for shipping a preview build to a small group of testers and getting back reports we can actually act on. Written 2026-09-21, the day the first playable build appeared.

Decisions already taken (owner's call):

  • Delivery: a "send report" button using the system share sheet. No backend, no background upload. The tester sees the file and sends it. This keeps us out of collecting data from other people's devices, which would otherwise need consent handling, storage and a retention policy.
  • Scope: crashes and bug reports, performance counters, unhonoured shim contracts, device model and OS version.

Why this needs building at all

Everything the engine currently reports goes to __android_log_write and nowhere else (util.cpp's Log()). That fails for beta in three separate ways, each of which already bit us during development:

  1. logcat wraps. Twice on 2026-09-21 a measurement was lost to it - once the whole startup sequence, once the frame-counter history, which made a frame-rate figure look far better than it was.
  2. Testers cannot retrieve logcat. It needs a cable and a developer setup.
  3. A hang produces no log at all. The black-screen bug that day left the process alive and silent; the only evidence was three lines that had already scrolled past.

Part 1 - what the build must produce

1.1 In-memory log ring (the foundation)

Log() keeps writing to logcat, and additionally appends into a fixed-size in-memory ring buffer (suggest 2 MB, tunable). Nothing is written to disk during normal play - the log rate is high (per-frame GLES sampling alone produces thousands of lines) and per-line file I/O would show up in the frame time we just spent the day reducing.

The ring is allocated once at startup. It must be usable from a signal handler, so: a plain pre-allocated byte array, no malloc, no locks that a crashed thread might hold. A per-record sequence number plus a short spinlock-free write is enough; a torn last record in a crash dump is acceptable.

1.2 Native crash handler

A handler for SIGSEGV, SIGBUS, SIGABRT, SIGILL, SIGFPE that writes a report file and then chains to the previous handler so the normal tombstone still happens.

Async-signal-safety is not optional here and we have a recorded incident: calling mmap() inside a SIGSEGV handler has deadlocked on bionic in this project before. The handler must use only a pre-allocated buffer, a pre-opened-or-open()ed fd, and write(). No malloc, no stdio, no Log(), no C++ allocation, no locks.

What to capture - this is the part specific to a translation layer, and it is what made today's crash solvable in one reproduction:

  • The decoded guest address. With the flat mapping, a host fault address is flat_map_base + guest_address. On 2026-09-21 the fault at 0x6f464c459b decoded to guest 0x464c459b by subtracting x28. The handler should do that subtraction itself and print the guest address, since nobody reading a report will do it by hand.
  • Which arena the guest address belongs to. guest_engine.cpp already has the classifier ("thread-stacks arena" and friends) - reuse it. "Fault in the heap arena" and "fault 1.2 GB past the end of everything" are completely different bugs and the report should say which.
  • Full guest register set, and the host registers from ucontext.
  • The last N KB of the log ring.
  • Build stamp and device identity (below).

1.3 Hang detector

The black-screen bug was a hang, not a crash - no signal, no tombstone, process alive at 2% CPU. A crash handler would have caught nothing.

A watchdog thread checks the onDrawFrame counter. If it has not advanced for ~10 seconds while the activity is resumed, it writes the same report the crash handler would, tagged HANG, including a snapshot of every thread's state and stack pointer. It should fire once per hang, not repeatedly.

1.4 Unhonoured-contract registry

Today's crash was found because a shim logged that it could not honour a request, one line above the fault. That should be a first-class, structured record rather than a log line we grep for.

A small registry: ReportUnhonouredContract(area, detail), deduplicated by string, counting occurrences. rtti_shims.cpp's use_facet, jni_shim.cpp's silent zero-returns (task #51) and every other "returning NULL because we do not implement this" path calls it. The report carries the full deduplicated list.

This turns beta into a gap-discovery mechanism: the union of these lists across testers is a prioritised work queue for what the game actually needs, discovered from real play instead of guessed.

1.5 Session counters

Sampled once a second into a compact rolling summary, not one line per sample:

  • frames per second - min, median, 10th percentile, and where the low ones happened (which is what the median alone hides, as it did today)
  • time from launch to first frame, and each level-load duration
  • CPU time consumed by the process, and by the GLThread specifically
  • peak CPU temperature
  • guest heap: live, peak, arena exhaustion events
  • thread-stack arena: peak in use, exhaustion events (the black-screen cause - this must never again be discovered by reading a log tail)

1.6 Build and device identity

Every report starts with: build stamp (version plus a short git hash or build timestamp baked in at compile time), device model, SoC, Android version, ABI, available RAM, and screen size.

The build stamp matters more than it sounds. Twice on 2026-09-21 the wrong APK was nearly measured - once a three-week-old release build installed by mistake. With testers there is no chance to check by hand; the report must say which build produced it.

1.7 The share button

A screen reachable from the pause menu: "Report a problem". It bundles the newest reports plus the current log ring into a single zip in the app's own files directory and hands it to ACTION_SEND.

Before sharing it shows the tester a short plain-language summary of what is in the file - log lines, device model, no personal data, no game account details. They are sending it themselves; they should know what it is.

If the previous session ended in a crash or hang, offer to send that report on the next launch, since the tester will not go looking for it.

Part 2 - the tester-facing report form

Free-text bug reports from testers are usually unusable not because testers are careless but because nobody told them which three facts matter. Keep it short - a long form gets skipped.

In-app, attached automatically: build stamp, device, the log bundle. The tester never types any of this.

What we ask the tester for, in this order:

  1. What were you doing? One line. "Entered a race from the city map."
  2. What happened? One line. "Black screen, music kept playing."
  3. What did you expect? Only when it is not obvious.
  4. Can you make it happen again? Every time / sometimes / happened once. This single question decides whether we can chase it at all.
  5. Did you play for a while before it happened? Yes/no. Specifically included because the whole class of resource-exhaustion bugs - the thread-stack arena, the guest heap - only shows up after a long session, and testers do not think to mention it.

Severity, defined by consequence rather than by feeling, so it is not argued about:

  • Blocker - cannot continue playing; progress lost.
  • Major - a feature does not work, but the session survives.
  • Minor - visual or audio defect, gameplay unaffected.

Ask them explicitly to send the report even when the game recovers. A hang that resolved itself still wrote a HANG report, and that is often the easier one to diagnose.

Part 3 - what we do with reports

Triage order, informed by what has actually been expensive to find:

  1. Unhonoured-contract list first, before reading the crash. Today the answer was in that list. It is cheap to check and frequently decisive.
  2. Decoded guest address and its arena. Distinguishes a wild pointer from arena exhaustion from a real logic bug, without any further work.
  3. Exhaustion counters. If a thread-stack or heap arena hit its ceiling, the crash is a symptom and the ceiling is the bug.
  4. Only then the register dump and the log tail.

Group reports by build stamp before comparing anything. Mixing builds is how a fixed bug looks like it is still present.

Suggested order of work

Each step is independently useful, so the beta does not wait on the whole set.

  1. Log ring + report file + share button, with device and build identity. Minimum shippable - a tester can send something useful.
  2. Crash handler with guest-address decoding.
  3. Unhonoured-contract registry, with use_facet and the JNI zero-returns as the first callers.
  4. Hang detector.
  5. Session counters.

Deliberately out of scope

  • No backend, no automatic upload. Chosen above. Revisit only if the manual path proves too lossy in practice.
  • No unique device or user identifier. Grouping by build stamp and device model is enough at this scale and avoids tracking individuals.
  • No gameplay telemetry - what cars, which races, how long played. It is not needed to fix defects, and collecting it would change what this file is.

Status 2026-09-21: the crash handler is built (plan section 1.2 + 1.7 partial)

Implemented and verified end to end on the Pixel 6a.

Native (mpcore/src/main/cpp/crash_handler.cpp). Hooks SIGSEGV, SIGBUS, SIGABRT, SIGILL, SIGFPE with SA_SIGINFO | SA_ONSTACK, on a pre-allocated alternate stack so a stack-overflow crash is still reportable. Everything the handler needs - the output path, the build stamp - is built at install time; inside the handler only open/write/close and hand-written integer formatters run. No malloc, no snprintf, no JNI. It chains to the previous handler afterwards, so Android still writes its own tombstone.

The report decodes the fault address: a host address inside the guest window is also printed as the guest address, and flagged when it is past the end of the mapped region ("a wild pointer, not a real guest object"). That is the number worth reading, and nobody will subtract the base by hand from a tester's report.

Java. CrashReportActivity renames the pending report (the handler writes a fixed name, since it cannot safely format a timestamp), zips it with device and build details, shows it, and offers ACTION_SEND through a FileProvider scoped to the crash directory only. Reports live in Android/data/<pkg>/files/crashes - reachable over USB with no permission.

Verified: handler installs; kill -11 produces a report with the right signal, registers and a correct "outside the guest window" verdict; the next launch detects it, renames it, builds the zip, and CrashReportActivity becomes the resumed activity. Files land where intended.

Not verified: what the screen actually looks like. The test device locked itself, so every screenshot was of a sleeping or locked display - which is also why an early "black screen" reading was wrong and led to one unnecessary fix (explicit colours, harmless and kept). The layout needs a human to unlock the phone and look.

Three failures on the way, all worth keeping

  1. Installed too early. The call sat at the top of onCreate, but libmpcore.so is only loaded later by loadCore() - UnsatisfiedLinkError, caught and logged. Moved to immediately after loadCore().
  2. Missing extern "C". The JNI function was C++-mangled (_Z65Java_...), so the JVM could not find it. The symptom was identical to the load-order bug above, which cost a wrong fix before the symbol table was actually read.
  3. Blocked activity start. Checking for a pending report in GameActivityMain.onCreate - which starts the report screen and finishes itself - was refused by the platform (BAL_ALLOW_GRACE_PERIOD) and dumped the tester on the home screen. The check belongs in PermissionsActivity, the visible launcher entry.

Status 2026-09-22: game data ships inside the APK

A tester now installs one file and plays. No separate .obb download, no file manager, no instructions about where to put anything.

How. The ~595 MB archive ships as assets/game_data.obb, and androidResources { noCompress += "obb" } keeps it stored rather than deflated - it is already compressed, so re-compressing would cost build and install time for nothing. On first launch GameDataUnpackActivity copies it to getObbDir()/main.<versionCode>.<package>.obb, which is exactly the path GameActivityMain.obbFullPath already builds, so no other code knows this happened.

The copy writes to a .part file and renames only on success. A half-written archive that merely exists would pass a naive check and send the game off to read truncated data - failing far from the cause, which is the failure mode this project keeps paying for. Free space is checked before starting rather than 500 MB in.

Measured on the Pixel 6a, with the real OBB renamed aside to simulate a clean device:

APK size 615 MB (was 22 MB)
build time 16 s - aapt2 handles the stored asset without trouble
adb install 31 s
unpack under 8 s - it finished before the first progress poll
result md5 identical to the original OBB
game afterwards 2,893 frames, 0 faults, mAssetLocationType=OBB

Costs worth stating. The device needs the APK plus the unpacked copy at once: about 1.2 GB free at install time, ~600 MB after. And the data exists twice on disk permanently, since Android keeps the APK.

The alternative not taken. Because the asset is stored uncompressed, its bytes sit contiguously in the APK, so the engine's own Shim_open/Shim_read could serve the OBB path straight out of the APK at an offset - no copy, no duplication. That is a real option if the 600 MB ever matters, but it adds a new failure surface in file I/O right before a beta, and the ask here was explicitly for self-extraction.