These files had never been tracked anywhere - they lived in a plain directory with no git at all, which is also where the whole reverse-engineering record sat. Code already committed refers to them by name (opponent_substitution.h cites "ANALYSIS.md section 6hh", DebugMenuOverlay.kt cites "DEBUG_MENU.md section 3"), so until now a fresh clone carried references to documents it did not contain. ANALYSIS.md the RE record, and the reason the rest works ARCHITECTURE.md how the mod's pieces fit together ARM64_TRANSLATION_LAYER.md the translation layer's running log PROGRESS.md chronological progress across both chats BETA_TELEMETRY_PLAN.md how crash/telemetry reporting is meant to work LOBBY_UI_DESIGN.md + .html lobby design and its clickable prototype DEBUG_MENU.md debug panel design STATIC_RECOMPILATION_FALLBACK.md the plan if translation had not panned out evidence/ font atlas capture from the glyph-corruption bug save_backups/ saves at known milestones, for reproducing state Co-Authored-By: Claude <noreply@anthropic.com>
14 KiB
Closed beta: diagnostics, crash reports and tester workflow
Plan for shipping a preview build to a small group of testers and getting back reports we can actually act on. Written 2026-09-21, the day the first playable build appeared.
Decisions already taken (owner's call):
- Delivery: a "send report" button using the system share sheet. No backend, no background upload. The tester sees the file and sends it. This keeps us out of collecting data from other people's devices, which would otherwise need consent handling, storage and a retention policy.
- Scope: crashes and bug reports, performance counters, unhonoured shim contracts, device model and OS version.
Why this needs building at all
Everything the engine currently reports goes to __android_log_write and
nowhere else (util.cpp's Log()). That fails for beta in three separate ways,
each of which already bit us during development:
- logcat wraps. Twice on 2026-09-21 a measurement was lost to it - once the whole startup sequence, once the frame-counter history, which made a frame-rate figure look far better than it was.
- Testers cannot retrieve logcat. It needs a cable and a developer setup.
- A hang produces no log at all. The black-screen bug that day left the process alive and silent; the only evidence was three lines that had already scrolled past.
Part 1 - what the build must produce
1.1 In-memory log ring (the foundation)
Log() keeps writing to logcat, and additionally appends into a fixed-size
in-memory ring buffer (suggest 2 MB, tunable). Nothing is written to disk
during normal play - the log rate is high (per-frame GLES sampling alone
produces thousands of lines) and per-line file I/O would show up in the frame
time we just spent the day reducing.
The ring is allocated once at startup. It must be usable from a signal handler,
so: a plain pre-allocated byte array, no malloc, no locks that a crashed
thread might hold. A per-record sequence number plus a short spinlock-free
write is enough; a torn last record in a crash dump is acceptable.
1.2 Native crash handler
A handler for SIGSEGV, SIGBUS, SIGABRT, SIGILL, SIGFPE that writes a
report file and then chains to the previous handler so the normal tombstone
still happens.
Async-signal-safety is not optional here and we have a recorded incident:
calling mmap() inside a SIGSEGV handler has deadlocked on bionic in this
project before. The handler must use only a pre-allocated buffer, a
pre-opened-or-open()ed fd, and write(). No malloc, no stdio, no Log(),
no C++ allocation, no locks.
What to capture - this is the part specific to a translation layer, and it is what made today's crash solvable in one reproduction:
- The decoded guest address. With the flat mapping, a host fault address is
flat_map_base + guest_address. On 2026-09-21 the fault at0x6f464c459bdecoded to guest0x464c459bby subtractingx28. The handler should do that subtraction itself and print the guest address, since nobody reading a report will do it by hand. - Which arena the guest address belongs to.
guest_engine.cppalready has the classifier ("thread-stacks arena"and friends) - reuse it. "Fault in the heap arena" and "fault 1.2 GB past the end of everything" are completely different bugs and the report should say which. - Full guest register set, and the host registers from
ucontext. - The last N KB of the log ring.
- Build stamp and device identity (below).
1.3 Hang detector
The black-screen bug was a hang, not a crash - no signal, no tombstone, process alive at 2% CPU. A crash handler would have caught nothing.
A watchdog thread checks the onDrawFrame counter. If it has not advanced for
~10 seconds while the activity is resumed, it writes the same report the crash
handler would, tagged HANG, including a snapshot of every thread's state and
stack pointer. It should fire once per hang, not repeatedly.
1.4 Unhonoured-contract registry
Today's crash was found because a shim logged that it could not honour a request, one line above the fault. That should be a first-class, structured record rather than a log line we grep for.
A small registry: ReportUnhonouredContract(area, detail), deduplicated by
string, counting occurrences. rtti_shims.cpp's use_facet,
jni_shim.cpp's silent zero-returns (task #51) and every other
"returning NULL because we do not implement this" path calls it. The report
carries the full deduplicated list.
This turns beta into a gap-discovery mechanism: the union of these lists across testers is a prioritised work queue for what the game actually needs, discovered from real play instead of guessed.
1.5 Session counters
Sampled once a second into a compact rolling summary, not one line per sample:
- frames per second - min, median, 10th percentile, and where the low ones happened (which is what the median alone hides, as it did today)
- time from launch to first frame, and each level-load duration
- CPU time consumed by the process, and by the GLThread specifically
- peak CPU temperature
- guest heap: live, peak, arena exhaustion events
- thread-stack arena: peak in use, exhaustion events (the black-screen cause - this must never again be discovered by reading a log tail)
1.6 Build and device identity
Every report starts with: build stamp (version plus a short git hash or build timestamp baked in at compile time), device model, SoC, Android version, ABI, available RAM, and screen size.
The build stamp matters more than it sounds. Twice on 2026-09-21 the wrong APK was nearly measured - once a three-week-old release build installed by mistake. With testers there is no chance to check by hand; the report must say which build produced it.
1.7 The share button
A screen reachable from the pause menu: "Report a problem". It bundles the
newest reports plus the current log ring into a single zip in the app's own
files directory and hands it to ACTION_SEND.
Before sharing it shows the tester a short plain-language summary of what is in the file - log lines, device model, no personal data, no game account details. They are sending it themselves; they should know what it is.
If the previous session ended in a crash or hang, offer to send that report on the next launch, since the tester will not go looking for it.
Part 2 - the tester-facing report form
Free-text bug reports from testers are usually unusable not because testers are careless but because nobody told them which three facts matter. Keep it short - a long form gets skipped.
In-app, attached automatically: build stamp, device, the log bundle. The tester never types any of this.
What we ask the tester for, in this order:
- What were you doing? One line. "Entered a race from the city map."
- What happened? One line. "Black screen, music kept playing."
- What did you expect? Only when it is not obvious.
- Can you make it happen again? Every time / sometimes / happened once. This single question decides whether we can chase it at all.
- Did you play for a while before it happened? Yes/no. Specifically included because the whole class of resource-exhaustion bugs - the thread-stack arena, the guest heap - only shows up after a long session, and testers do not think to mention it.
Severity, defined by consequence rather than by feeling, so it is not argued about:
- Blocker - cannot continue playing; progress lost.
- Major - a feature does not work, but the session survives.
- Minor - visual or audio defect, gameplay unaffected.
Ask them explicitly to send the report even when the game recovers. A hang
that resolved itself still wrote a HANG report, and that is often the easier
one to diagnose.
Part 3 - what we do with reports
Triage order, informed by what has actually been expensive to find:
- Unhonoured-contract list first, before reading the crash. Today the answer was in that list. It is cheap to check and frequently decisive.
- Decoded guest address and its arena. Distinguishes a wild pointer from arena exhaustion from a real logic bug, without any further work.
- Exhaustion counters. If a thread-stack or heap arena hit its ceiling, the crash is a symptom and the ceiling is the bug.
- Only then the register dump and the log tail.
Group reports by build stamp before comparing anything. Mixing builds is how a fixed bug looks like it is still present.
Suggested order of work
Each step is independently useful, so the beta does not wait on the whole set.
- Log ring + report file + share button, with device and build identity. Minimum shippable - a tester can send something useful.
- Crash handler with guest-address decoding.
- Unhonoured-contract registry, with
use_facetand the JNI zero-returns as the first callers. - Hang detector.
- Session counters.
Deliberately out of scope
- No backend, no automatic upload. Chosen above. Revisit only if the manual path proves too lossy in practice.
- No unique device or user identifier. Grouping by build stamp and device model is enough at this scale and avoids tracking individuals.
- No gameplay telemetry - what cars, which races, how long played. It is not needed to fix defects, and collecting it would change what this file is.
Status 2026-09-21: the crash handler is built (plan section 1.2 + 1.7 partial)
Implemented and verified end to end on the Pixel 6a.
Native (mpcore/src/main/cpp/crash_handler.cpp). Hooks SIGSEGV, SIGBUS,
SIGABRT, SIGILL, SIGFPE with SA_SIGINFO | SA_ONSTACK, on a pre-allocated
alternate stack so a stack-overflow crash is still reportable. Everything the
handler needs - the output path, the build stamp - is built at install time;
inside the handler only open/write/close and hand-written integer
formatters run. No malloc, no snprintf, no JNI. It chains to the previous
handler afterwards, so Android still writes its own tombstone.
The report decodes the fault address: a host address inside the guest window is also printed as the guest address, and flagged when it is past the end of the mapped region ("a wild pointer, not a real guest object"). That is the number worth reading, and nobody will subtract the base by hand from a tester's report.
Java. CrashReportActivity renames the pending report (the handler writes a
fixed name, since it cannot safely format a timestamp), zips it with device and
build details, shows it, and offers ACTION_SEND through a FileProvider scoped to
the crash directory only. Reports live in
Android/data/<pkg>/files/crashes - reachable over USB with no permission.
Verified: handler installs; kill -11 produces a report with the right
signal, registers and a correct "outside the guest window" verdict; the next
launch detects it, renames it, builds the zip, and CrashReportActivity becomes
the resumed activity. Files land where intended.
Not verified: what the screen actually looks like. The test device locked itself, so every screenshot was of a sleeping or locked display - which is also why an early "black screen" reading was wrong and led to one unnecessary fix (explicit colours, harmless and kept). The layout needs a human to unlock the phone and look.
Three failures on the way, all worth keeping
- Installed too early. The call sat at the top of
onCreate, butlibmpcore.sois only loaded later byloadCore()-UnsatisfiedLinkError, caught and logged. Moved to immediately afterloadCore(). - Missing
extern "C". The JNI function was C++-mangled (_Z65Java_...), so the JVM could not find it. The symptom was identical to the load-order bug above, which cost a wrong fix before the symbol table was actually read. - Blocked activity start. Checking for a pending report in
GameActivityMain.onCreate- which starts the report screen and finishes itself - was refused by the platform (BAL_ALLOW_GRACE_PERIOD) and dumped the tester on the home screen. The check belongs inPermissionsActivity, the visible launcher entry.
Status 2026-09-22: game data ships inside the APK
A tester now installs one file and plays. No separate .obb download, no file manager, no instructions about where to put anything.
How. The ~595 MB archive ships as assets/game_data.obb, and
androidResources { noCompress += "obb" } keeps it stored rather than
deflated - it is already compressed, so re-compressing would cost build and
install time for nothing. On first launch GameDataUnpackActivity copies it to
getObbDir()/main.<versionCode>.<package>.obb, which is exactly the path
GameActivityMain.obbFullPath already builds, so no other code knows this
happened.
The copy writes to a .part file and renames only on success. A half-written
archive that merely exists would pass a naive check and send the game off to
read truncated data - failing far from the cause, which is the failure mode this
project keeps paying for. Free space is checked before starting rather than
500 MB in.
Measured on the Pixel 6a, with the real OBB renamed aside to simulate a clean device:
| APK size | 615 MB (was 22 MB) |
| build time | 16 s - aapt2 handles the stored asset without trouble |
adb install |
31 s |
| unpack | under 8 s - it finished before the first progress poll |
| result | md5 identical to the original OBB |
| game afterwards | 2,893 frames, 0 faults, mAssetLocationType=OBB |
Costs worth stating. The device needs the APK plus the unpacked copy at once: about 1.2 GB free at install time, ~600 MB after. And the data exists twice on disk permanently, since Android keeps the APK.
The alternative not taken. Because the asset is stored uncompressed, its
bytes sit contiguously in the APK, so the engine's own Shim_open/Shim_read
could serve the OBB path straight out of the APK at an offset - no copy, no
duplication. That is a real option if the 600 MB ever matters, but it adds a new
failure surface in file I/O right before a beta, and the ask here was explicitly
for self-extraction.