Commit Graph

25 Commits

Author SHA1 Message Date
Kgothatso Ngako
2309879153 test(frost): cover the batch's failure modes and its crypto without a database
Phase 6 of docs/frost-batch-signing.md. 361 jvmTest and 227 testDebugUnitTest
pass.

## Inbound path (SignedGroupKeyStateTest)

Both drive the manager with a hand-built inner event rather than one the other
device queued, which is the only way to be a faulty or dishonest member in this
harness.

- A one-value nonce offered for a three-item batch does not count towards the
  threshold: the coordinator never reaches a signer set. The length check is all
  that stands between a batch and a signer whose contribution lines up against
  the wrong messages, so truncating or padding would produce partial signatures
  aggregated against events nobody agreed to. The test then pumps the real nonce
  and the batch completes -- it is a stall, not damage, which is
  FrostSignerMessage's composite key doing its job.
- A second proposal under the session's own id changes neither its event ids nor
  its seeds. Every seed is already committed to its item's message; a different
  batch under the same id would have those seeds produce a second partial
  signature over a second message, which is how a share is extracted.

## Real FROST, no database (FrostSigningRoundTest)

- A k=3 batch from one signer set, all three verifying against the room's key --
  the manager's shape with the database taken out of the way.
- Item 0's signature does not verify against item 1. Signing three events in
  lockstep must not make any of them interchangeable.
- Both halves of the no-shared-nonce property, because either alone is enough to
  be relied on by accident: SecretNonce.generate mixes the message in, so one
  seed under two messages already gives two nonces -- and the manager mints
  distinct seeds regardless.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-06 04:58:48 +02:00
Kgothatso Ngako
59c34263b3 feat(frost): let one signing session carry a batch of events
Phase 3 of docs/frost-batch-signing.md. A session can now be proposed over
several events, and the whole batch is signed in one round of four group
events with one approval. 356 jvmTest and 224 testDebugUnitTest pass.

## The wire, and the compatibility rule that shapes it

FrostSigningEvents.encodeProposal serialises a batch of one as the bare event
object it always was, and only a genuine batch as a JSON array. That is not
tidiness. A build predating this reads an array with Event.fromJsonOrNull, gets
null, and drops the proposal -- so an old device refuses a batch outright
rather than signing part of one, while single signing keeps working right
through a mixed-version rollout. Emitting an array unconditionally would break
every one-event session for those devices and buy nothing.

decodeProposal accepts both forms permanently: proposals in the old shape do
not stop arriving because this build stopped writing them. It is
all-or-nothing -- an array with one unreadable element is refused rather than
silently shortened, because the batch's length is what every later payload is
checked against, and a proposal that quietly lost an event would have every
signer's contribution rejected for being the wrong size: a stall with nothing
to blame.

## MAX_BATCH_SIZE, checked twice

64, enforced in proposeSigningBatch and again, independently, in
acceptProposal. The second check is the one that matters. A proposal is the
only place in this protocol where a remote party decides how much work everyone
else does -- k native key generations, k signatures, and a group event carrying
k payloads, from a single message -- and until batching that was bounded only
by never being more than one.

## acceptProposal over a list

Each element is rebuilt from its own fields under this device's own reading of
the room's path and checked against the id it claims, exactly as before but per
item, and the whole proposal is dropped if any one fails. The write-once rule
widens from "the event this session signs" to "the ordered list of events this
session signs": a second proposal under the same id whose list differs anywhere
is logged and ignored.

## The API

FrostSigningManager.proposeSigningBatch(events: List<EventTemplate<*>>) is
public here rather than in Phase 4, because without it there is no way to
produce a k>1 session and everything above would ship untested. proposeSigning
keeps its signature as the one-event form, so no caller moves. Each template
carries its own createdAt.

## Tests

- FrostProposalCodecTest (new, commonTest): a batch of one is byte-for-byte the
  old JSON object -- the assertion that stands in for the old build nobody can
  run here -- plus order preservation, old-form decoding, and refusal of empty,
  malformed and partly-unreadable arrays.
- SignedGroupKeyStateTest: a k=3 batch between two devices over two databases.
  Three signatures verifying against the room, three dialects applied on both
  devices in order, five messages from the coordinator and two from the other
  signer, and one approval line rather than three.
- The negative test that matters: no two items of a batch share an aggregated
  nonce or a seed, and the two devices' seeds do not intersect. Every positive
  test still passes if two items share a nonce -- the signatures verify fine;
  what sharing costs is the secret share.
- The cap is refused when proposed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-06 04:48:24 +02:00
Kgothatso Ngako
935a8fe37a refactor(frost): run a signing session as k FROST instances in lockstep
Phase 2 of docs/frost-batch-signing.md. Pure refactor: proposals still carry
one event, the wire is byte-identical, and every test passes unchanged --
344 jvmTest and 217 testDebugUnitTest, none of them edited in this commit.

advance() now loops over FrostSigningItem rows rather than reading the first
one. One nonce per item, one aggregate per item, one Session.create per item,
one partial signature per item, one signature per item. The signer set, the
public shares, the tweak cache and the approval stay shared, because they are
the terms that do not enter e = H(R‖P‖m).

The coordinator's aggregation is the place where that distinction bites: it
builds one AggregatedNonce per item, each from that item's nonce from each
chosen signer. Reusing one across two items would be reusing R across two
messages.

## The payload codec, early

joinPayload/splitPayload land here rather than with the wire change, because at
a batch of one a comma join is the identity -- the payload is the bare value it
has always been. That leaves Phase 3 to the proposal encoding alone.

splitPayload is strict: a payload that is not exactly the batch's length is
dropped rather than truncated or padded. It runs in orderedNonces,
orderedPartialSignatures and splitForSession -- never in record(), which stores
payloads without parsing them so that a nonce can arrive before the proposal
that would give it a length to check against.

## Two short-circuits, and one trap in the first

advance() runs on every arriving message, so at a batch of k it was k native
key generations, k Session.creates and k signs each time, usually to discover
there was nothing left to do.

- Nonces are generated by `lazy`. The obvious version -- a guard computing
  `ownNonce == null || (isSigner() && ownPartial == null)` -- is wrong, and
  wrong in a way that reads fine and fails every signing test: the coordinator
  settles the signer set further down the same pass, so isSigner() at the top
  is false on exactly the pass where the coordinator goes on to sign, and the
  nonces are never generated. Reproduced as IndexOutOfBounds before switching
  to lazy, which has no prediction to make.
- A device that is neither signing nor aggregating leaves before building any
  FROST session, rather than building k of them to do nothing with.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-06 04:40:22 +02:00
Kgothatso Ngako
dff41d417d feat(frost): move a signing session's per-event columns onto FrostSigningItem
Phase 1 of docs/frost-batch-signing.md, which is added here as the plan the
next phases follow. Schema only: a session still signs exactly one event, the
wire is byte-identical, and every existing test passes on the moved columns.

## What moved, and why it had to

A batch of k events is k independent FROST instances sharing a signer set, not
one signature over k messages. That is forced rather than chosen: a Schnorr
partial signature is `s = k + e·x` with `e = H(R‖P‖m)`, so two messages under
one nonce R give two equations in one unknown and the secret share falls out.

So the five columns that enter that equation -- unsignedEventJson, eventId,
nonceRandom, aggregatedNonce, signature -- move to a child table keyed
(sessionId, itemIndex). What stays on FrostSigningSession is everything outside
it: the ceremony, the threshold, the derivation path, the signer set, and the
one approval.

itemIndex is protocol rather than presentation -- nonces and partial signatures
are joined positionally against it -- so getItems() orders by it and nothing
re-sorts. Spelled itemIndex rather than index to keep hand-written queries free
of backticks.

No itemCount column. The count is a COUNT(*), for the same reason signerIds is
derived from the ceremony's participant order rather than stored: a
denormalised count is one more thing that can disagree with the rows.

## Migration 9 -> 10

Manual, not auto: Room can create the table and drop the columns but cannot
copy between them, and the copy is the whole point. A session in flight at
upgrade holds its nonce seed and the aggregate it is already signing against,
and neither can be regenerated -- losing either makes the next pass derive a
different nonce for the same message and publish a second partial signature
over it, which is the extraction case. Both are copied verbatim into item 0, so
an in-flight session resumes as though nothing happened.

Removing the columns uses ALTER TABLE DROP COLUMN rather than the usual
create-copy-drop-rename rebuild. FrostSignerMessage and FrostSigningItem both
reference FrostSigningSession(id) ON DELETE CASCADE, and DROP TABLE fires
cascades -- with foreign keys enforced the rebuild would delete every signer
message and every item just written. Whether it does depends on Room disabling
foreign keys around migrations, which is not worth depending on when
DROP COLUMN cannot go wrong. It needs SQLite 3.35 and unindexed,
unconstrained columns; these five qualify, and getRoomDatabase pins
BundledSQLiteDriver on every platform.

## Invariants established here for the phases that follow

- signerIds and every item's aggregatedNonce are one write-once unit, applied
  by applyAggregate() -- items first in one transaction, then the session, so
  "some items aggregated" is unreachable and signerIds != null stays the gate.
- Signatures likewise, via applySignatures(); isSigned() counts rows instead of
  reading a flag.
- complete() verifies every signature before applying any event, so a batch is
  all-or-nothing rather than half-filed.
- itemsOver() gives each item its own 32 bytes of seed. Independent seeds mean
  an off-by-one in index handling produces a session that fails to aggregate
  rather than one that signs two messages under a single nonce.

signedEvent() and isAwaitingApproval() now take the item(s) rather than the
session, which propagates to the repository, the view model and the screen.
advance() reads items.first() and Phase 2 turns that into a loop.

## Tests

- FrostSigningSessionDaoJvmTest: index ordering, single-item read, upsert
  replacing rather than accumulating, signed-item counting, cascade delete.
- FrostSigningItemMigrationJvmTest (new): the backfill against a real v9
  database, asserting the seed and aggregate values survive -- not merely that
  a row appeared -- plus the exact column lists Room will check at open time.
- 338 jvmTest and 217 testDebugUnitTest pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-06 04:33:46 +02:00
Kgothatso Ngako
af81933ab4 Merge branch 'mantra' into claude/room-db-testing-setup-b053cd
Brings the branch up to date with the 40 commits mantra gained while the
jvm target was being built, so that merging the other way is a
fast-forward.

One conflict, in docs/README.md, where both sides added rows to the index
table. Kept both, and gave the jvm-target note a clause in the closing
prose since it is the one document there that is not about the protocol.

One thing the auto-merge could not have caught. `9250991` added
NostrEventDao.getMarmotGroupNostrEventsByChatRoomId as a blocking query,
which android accepts and which Room refuses to generate for any other
target -- so the merged tree failed :composeApp:compileKotlinJvm with the
same "Only suspend functions are allowed in DAOs declared in source sets
targeting non-Android platforms" that phase 4 dealt with 58 times. Made
suspend; its only caller, NostrDao.reindexMarmotGroupEvents, was already
suspend, so again no cascade.

That is now a standing cost of this branch rather than a one-off: any DAO
method added on mantra while this is outstanding will break the jvm build
on merge. It is a one-word fix each time, and the compiler names the line.

Verified on the merged tree: :composeApp:compileKotlinJvm and
:composeApp:compileDebugKotlinAndroid green,
:composeApp:testDebugUnitTest 208 passing, :composeApp:jvmTest 214
passing -- both test tasks re-run from scratch rather than taken from the
cache.

The jvm figure is larger than the android one because jvmTest inherits
commonTest, so declaring the target quietly gained the whole shared suite
a second execution environment. That is worth knowing independently of
whether desktop ever ships: the same tests now run on the host, without an
emulator.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-06 02:01:35 +02:00
Kgothatso Ngako
adf1f03817 feat: a desktop entry point, and the first code here that runs
Phase 5. `press.mantra.desktop.MainKt` has been named by the
compose.desktop block since before this work started and did not exist;
now it does, and `./gradlew :composeApp:run` opens a window.

**The window opens onto a passphrase gate, not onto the app.** That is
phase 3 landing here rather than there, and it was not in the plan.
keyStoreEncryption(keyName, plainText) takes no secret, because on android
the OS keystore serves keys without asking anybody anything -- so a
passphrase scheme needs an unlock the expect signature cannot express.
MainKt calls JvmKeyStore.unlock before MantraApp is composed, off the ui
thread, because Argon2id at 64 MiB is deliberately slow enough to stop the
window painting. The gate says on its face that this build is not for real
funds.

One application directory is handed to both the mantra and the phoenix
context, so a single install keeps a single place on disk rather than two
named after different projects.

**MantraDatabaseJvmTest is the part worth keeping.** Running the app
proves the window paints; it proves nothing about Room, because the gate
stops before anything touches the database. Six tests now open it: the
schema is created, a profile survives a write and a read, upsert replaces
rather than duplicates, the @Transaction relation query behind
findChatRoomById reads back, a soft-deleted room stops being found, and
the on-disk builder writes under the context directory rather than
java.io.tmpdir. This is the first time this database has been opened
anywhere but android, and it covers exactly what the compiler cannot see
-- that Room's ksp output for this target is usable, that the *host*
SQLite native loads where the android artifact's would not, and that the
58 queries forced from blocking to suspend still return what they stored.

Both of that test's first drafts were wrong in ways worth keeping the
scars of. Every write failed with SQLite error 787 because Profile has a
foreign key onto NostrEvent and the test never created the parent row --
which is evidence rather than an annoyance, since a schema whose
constraints were quietly off would have let all of it pass. And Kind is a
typealias for Int, not a constructor.

Window sizing is 480x900: a starting size that does not immediately
misrepresent layouts only ever exercised at phone widths, not a considered
desktop layout. That, along with back handling and any ui offering an nfc
affordance, is the shakeout this phase names and does not do.

Verified, all five green: :composeApp:compileKotlinJvm,
:composeApp:compileDebugKotlinAndroid, :composeApp:testDebugUnitTest (52),
:composeApp:jvmTest (6), and the fork's :library:jvmTest (97).

Not verified: nothing past the gate. No seed has been written, no business
started, no relay contacted. A gradle `run` killed with SIGTERM reports
BUILD FAILED with exit value 143 -- that is the signal, not the app.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-06 01:34:10 +02:00
Kgothatso Ngako
bf4041a5b2 feat: mantra compiles for the jvm
Phase 4. Declares jvm(), implements all 16 expects, and bumps the
submodule to the fork branch carrying phases 1-3.
:composeApp:compileKotlinJvm is green.

**The actuals were the small half. Room was the blocker.** The first jvm
compile failed with 58 copies of "Only suspend functions are allowed in
DAOs declared in source sets targeting non-Android platforms". Room
permits blocking query methods on android and nowhere else, so every @Dao
function that was neither suspend nor Flow-returning had to change --
58 of them across 24 files. KSP reports these in alphabetical batches, so
the count shrinks in stages and looks bottomless; scanning the dao package
directly for abstract funs with no suspend and no Flow return finds them
all at once.

It stops there, which is the only reason this is a 58-line change rather
than a refactor. Every one of the 15 call sites outside the dao package
was already inside a suspend function -- the repositories were written
that way throughout -- so nothing needed rewriting. One private helper,
DatabaseNostrRepository.matchNegentropicNostrEvents, had to become
suspend, and its single caller was already suspend, so the cascade
terminated immediately. Zero call-site edits.

**The cost lands on android, not on the jvm.** A blocking DAO method runs
on its caller's thread; a suspend one is dispatched to the query coroutine
context, which getRoomDatabase sets to Dispatchers.IO. That is the better
behaviour -- it is what stops a query running on the main thread -- but it
is a real change to the shipping platform, made for a target that does not
run yet. Hence the unit tests below rather than a compile alone.

**BusinessManager was not an expect**, so nothing warned about it. It is
now ported to the fork's jvmMain (05ce7eb); Phoenix.jvm.kt and
NavigationViewModel.jvm.kt are otherwise the ios actuals with one changed
import, since those files use no ios API.

**schedulePlatformLogic schedules nothing, and logs that it does not.**
Android starts two WorkManager jobs here, one of which is ChannelsWatcher
-- it wakes periodically to notice a channel force-closed while the app
was shut. A desktop application has no process once its window closes, so
there is nothing to wake, and running the watcher in-process would be
strictly worse than not running it: it would only fire while the app was
already open and watching. The exposure is real and belongs in release
notes rather than a comment -- a desktop wallet left closed past a
force-close deadline does not notice.

Smaller calls. PlatformContext carries an application directory, since
there is no Context to read one from, and PlatformDatabaseBuilder puts
aux.db under it rather than in java.io.tmpdir, which is what the abandoned
Aux implementation did behind a TODO and which most systems clear on
reboot. themeColorScheme ignores dynamicColor, which means Material You
and has no desktop counterpart. AppVersion reads the jar manifest that
compose.desktop writes, falling back when running from a class directory.

Verified: :composeApp:compileKotlinJvm green,
:composeApp:compileDebugKotlinAndroid green, and
:composeApp:testDebugUnitTest 52 passing -- the one that matters, since
this commit changes shared code every android query path goes through.

Not verified: nothing has run. No jvm entry point exists yet, so the
database has never been opened on this platform and no business has been
started. That is phase 5, which also has to unlock JvmKeyStore before the
wallet starts -- a passphrase prompt, not just a window.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-06 01:23:42 +02:00
Kgothatso Ngako
670f87a609 docs: record what phase 3 turned out to require
Phase 3 is implemented in the fork on claude/jvm-target-actuals (ce49657).
The security analysis in the plan held up; three practical constraints
around it did not appear until the code was written.

**A passphrase-derived KEK is not a drop-in.** The plan treated the choice
between a passphrase, an OS keychain and a key file as the whole decision.
But keyStoreEncryption(keyName, plainText) takes no context and no secret
-- on android the OS holds the key, so none is needed -- which means any
passphrase scheme needs an out-of-band unlock the expect cannot express.
That is a change to application startup, not just to the actual, so it is
now called out against phase 5: the desktop entry point has to prompt and
unlock before the wallet starts.

**The iv must be 16 bytes.** EncryptedSeed.V2.serialize in commonMain
throws on anything else, which rules out a conventional 96-bit GCM nonce
-- worth knowing before designing around one. It turns out to help: with
randomly generated nonces the risk is a repeat under one key, and 128 bits
makes that vanishingly unlikely where 96 merely makes it unlikely.

**Argon2id costs a dependency.** The jdk has PBKDF2 and no memory-hard
KDF at all, so it means bouncycastle. Recorded with the reason to pay it:
if the build is dev-only because it lacks hardware backing, weakening the
KDF too gets the trade backwards.

Also recorded: wrap a per-name data key under the KEK rather than
encrypting the seed with it directly, so a passphrase change rewraps 32
bytes; throw java.security.KeyStoreException when locked, since the
graceful* wrappers already map it to
DecryptSeedResult.Failure.KeyStoreFailure; and a verification section
naming the properties that fail quietly, plus the two limits worth writing
down rather than fixing -- the first unlock on a new store accepts any
passphrase, and zeroing the key is best effort.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-06 01:13:07 +02:00
Kgothatso Ngako
506dfea802 docs: correct phase 2 against what the jdbc drivers actually needed
Phase 2 is implemented and verified in the fork on
claude/jvm-target-actuals (d1a82ea). Three corrections and one omission.

**Schema handling does not need hand-rolling.** The plan said "you call
Schema.create(driver) and the migration path explicitly, and you have to
track the applied version yourself". SQLDelight 2.x ships a factory
function that shadows the constructor -- JdbcSqliteDriver(url, properties,
schema, migrateEmptySchema, vararg callbacks) -- which does all three,
user_version included. The same-named constructor does none of it, which
is the trap worth naming rather than the work that was budgeted for.

**Foreign keys were the actual work, and the plan never mentioned them.**
Off by default in SQLite, and the pragma is per connection while
JdbcSqliteDriver opens one per thread, so it has to go through the
connection Properties rather than be issued once against the driver.
Recorded along with why that needs no compile dependency on
org.xerial:sqlite-jdbc, which arrives at runtime scope only.

**commonTest has an expect too.** The 23 counted at the top of this
document are commonMain's. Declaring jvm() also creates jvmTest, which
inherits commonTest, so `connect` in ElectrumServersTest blocks every jvm
test from compiling. Noted along with the reason not to stub it empty the
way ios does: the class is @Ignore'd everywhere, so an empty body looks
harmless right up until somebody removes the @Ignore and
connect_to_mainnet_servers starts passing without connecting to anything.

**Phase 2 is the first phase that can be run, and the plan told you not to
bother.** It said "none of this is exercisable until Phase 4. Write the
SQLDelight schema-creation path against a scratch main() if you want
feedback sooner." That was wrong twice: library/src/jvmTest/ already
exists, and the two properties worth checking are exactly the ones a
compiler cannot see. Schema creation and the foreign-key pragma both fail
silently in production -- a missing table only shows up at first query,
and foreign keys being off means cascading deletes quietly do not happen.
The phase now carries a real exit condition, and DbFactoryJvmTest meets it
with five passing tests.

Recorded with it: the two KeyStoreFunctions actuals have to exist before
phase 3 decides anything, because nothing jvm compiles without them, and
they should throw rather than do something plausible.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-06 01:04:02 +02:00
Kgothatso Ngako
0757e50dc5 docs: correct the jvm plan against what phase 1 actually did
Phase 1 is implemented and verified in the lightning-kmp-app fork on
claude/jvm-target-actuals (27a0054). Four things in the plan were wrong,
and doing the work is what surfaced them.

**jvm() belongs at the start of phase 1, not phase 4 -- for the library.**
The plan said leave it off in both builds until phase 4. That is right for
mantra and wrong for the fork: library/src/jvmMain/ is an orphan source
set until the library declares the target, so phases 1-3 would all have
been written blind. Declared first, `:library:compileKotlinJvm` names the
remaining expects, and that list beats grepping for `expect ` -- it
shrinks by exactly what you implement and cannot drift from the truth. The
build stays red across phases 1-3 by design.

That checklist is now recorded as the phase 1 exit condition: exactly
eight expects should remain, and exactly which eight. Anything else means
something in the phase is wrong.

**Phase 3 is two decisions, not four.** gracefulSingleSeedDecryption and
gracefulMultiSeedDecryption are pure exception mapping into a
DecryptSeedResult, and the exception they branch on is
java.security.KeyStoreException -- a plain JCA type that exists on the jvm.
Both are near-copies of the android actuals and need nothing settled
first, so they move alongside phase 2. Only keyStoreEncryption and
keyStoreDecryption are the security decision, and that part of the
analysis stands.

**The Fibonacci template must not be deleted.** The plan said to drop it
"assuming nothing references them". Things do: generateFibi is exercised
by template tests in commonTest, androidHostTest, iosTest, jvmTest and
linuxX64Test, and JvmFibiTest asserts a value that depends on precisely
the two properties fibiprops.jvm.kt defines. That file already satisfies
two of the 25 expects, which is why the count was 23 missing rather than
25. Removing the template is five test files plus four fibiprops.*
actuals, and it is a separate cleanup.

**Phase 1 is fifteen actuals, not fourteen**, and two of them are not
copies of android -- platformElectrumRegtestConf (10.0.2.2 is the
emulator's alias for the host loopback; a jvm process is already on the
host) and phoenixLogWriters (android routes kermit into slf4j because
android tooling reads that back).

Also recorded, because it cost time: a worktree cannot run gradle at all
until the submodules are checked out *and* local.properties exists at five
levels. Neither is version controlled, so a fresh worktree has neither,
and the failure surfaces four builds down at
:...:secp256k1-kmp:jni:android as "SDK location not found" rather than
anywhere obviously related.

Both builds were run: `:library:compileKotlinJvm` fails only on the known
eight, and `:composeApp:compileDebugKotlinAndroid` still passes with the
library's jvm target declared -- the check that matters, since a new
variant must not change how the android target resolves the library.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-06 00:53:51 +02:00
Kgothatso Ngako
2aaa7b99a6 build: phase 0 of the jvm target -- clear the ground, correct the plan
First phase of docs/jvm-target.md. Nothing here turns the target on; it
removes what would break the moment it is turned on, and stages the two
catalog entries that cannot be derived automatically. Two of the four
steps as written in the doc turned out to be wrong, and implementing them
is how that surfaced -- both are corrected in the doc in this commit.

**Deleted the stale jvmMain tree.** Six files under
composeApp/src/jvmMain/kotlin/ac/cord/auxiliary/ survived from the Aux
project this codebase grew out of. They have gone unnoticed because
`jvmMain` is currently an orphan source set -- the accessor creates it,
no target compiles it -- so the wrong package, the Room 2 imports
(androidx.room, not androidx.room3), and the references to a long-gone
AuxDatabase and AuxGlobal have never had to resolve. They would all become
compile errors in phase 4.

They are not lost: they are the closest thing to a skeleton for five of
the six platform actuals phase 4 needs, and main.kt is a reasonable
starting shape for the phase 5 desktop entry point. `git show HEAD~1` has
them.

**Added two catalog entries, not four.** sqlite-bundled-jvm and
sqldelight-sqlite-driver. Both earn their place by being unreachable
otherwise: sqlite-bundled-jvm has to be named explicitly because
variant-aware resolution hands the *android* artifact to anything running
on the host, and sqldelight-sqlite-driver is the jvm counterpart to the
android-driver and native-driver entries already there.

The doc also listed room3-runtime-jvm and sqldelight-jdbc-driver. Neither
is right. Once jvm() exists, commonMain's existing androidx-room3-runtime
resolves to the -jvm variant on its own, so an explicit entry is
redundant and would drift. And the SQLDelight drivers phase 2 needs are
for DbFactory, which lives in lightning-kmp-app -- a separate gradle build
with its own version catalog, where an entry here is simply not visible.

**kspJvm cannot be wired yet, and the build file already said so.** The
doc's phase 0 told you to uncomment
composeApp/build.gradle.kts:194. It contradicted its own phase 4, which is
where jvm() gets turned on. The comment three lines above it states the
rule:

    These configurations only exist when the ios targets are declared,
    which the kotlin block above does only on a mac.

The same holds for kspJvm -- `dependencies { add("kspJvm", ...) }` throws
UnknownConfigurationException until a jvm() target creates the
configuration. So it moves into phase 4, into the same edit that declares
the target. composeApp/build.gradle.kts is deliberately untouched by this
commit.

**Also documented: gradle does not run in a worktree here at all** until
the submodule is checked out, which worktrees do not do automatically.
`lightning-kmp-app/` is empty and configuration fails with "Project with
path ':library' not found in build ':lightning-kmp-app'". Recorded in the
phase 0 verification section along with the caveat that a linked worktree
shares .git/modules/ with the main checkout, so both trees end up on one
submodule git dir.

**Not verified by a build.** For that reason. The deletion is an orphan
source set and the additions are unreferenced catalog lines, so neither
can change a build's outcome -- but that is an argument, not a green
check, and it is the second commit in a row on this branch that has not
compiled anything. Phase 4 is the first phase that genuinely cannot be
done without a working gradle invocation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-06 00:44:10 +02:00
Kgothatso Ngako
65e4a3acc0 fix: seal the Welcome, the one gift wrap an MLS room must publish
No invite to a Marmot room has been delivered since 1700e6d. The Welcome
was built, hashed, queued and logged exactly as before -- and then refused
one step short of the relay, by a guard that had no idea it was looking at
one.

**The guard.** 1700e6d ("send a direct message into the group, wrapped for
one member") added a backstop at the top of sealGiftWrapPayload:

    val chatRoom = database.chatRoomDao().findChatRoomById(giftWrapPayload.chatRoomId)
    if (chatRoom?.chatRoom?.mlsGroupState != null) { log; return }

Its reasoning is sound and still is: a Marmot direct message is a genuine,
correctly signed NIP-59 gift wrap, indistinguishable from one this path
would be right to publish, and the only thing keeping it off a relay is
that it never becomes a GiftWrapPayload row. Refusing at the seal as well
means a future caller cannot walk one onto a relay by accident.

**Why it caught the Welcome.** MarmotOutboundDao.deliveryWelcome writes a
GiftWrapPayload with chatRoomId = nostrGroupId -- the MLS room's own id,
which by construction has mlsGroupState set. Every Welcome therefore
matched a refusal keyed on mlsGroupState alone. There is no Welcome that
does not: the tag identifying the room is the whole point of the event.

The kind:444 arm further down -- the one that wraps only for the
participant who published the referenced key package -- became
unreachable, which is why nothing in the logs said "welcome" at all.

**Why it went unnoticed.** The room reaches the correct state on the
inviter's side whether or not the Welcome goes out: addMember advances the
epoch, the group state is persisted, the Participant row exists, and
"Invited X to chat" is written to the transcript. From the coordinator's
side an invitee who never heard anything is indistinguishable from one who
joined -- recorded as a known gap in docs/shared-key-ceremony.md, and this
is what was behind it.

**The second-order damage.** observeUnsealedGiftWrapPayloads is

    SELECT * FROM GiftWrapPayload WHERE publicKey = :p AND giftWrapSealId IS NULL

collected as a Flow<GiftWrapPayload?> -- one row at a time. giftWrapSealId
is only ever set inside persistAndBroadcastGiftWrap, which the guard
returns before reaching, so a refused payload stays unsealed forever and
sits at the head of that queue. The first Welcome a user queued blocked
every gift wrap behind it, in every room, for the life of the install.

**The fix.** Kind first, room second. MIP-02 addresses kind:444 to someone
who is not yet in the group and holds no key to read a kind:445 -- a
relay-borne gift wrap is the only way to reach them, and deliveryWelcome
queues one on purpose. Every other kind is refused exactly as before.
sendChatMessage already branches on mlsGroupState before writing a
payload, so an MLS room's messages never arrive here anyway; the guard
stays as the backstop it was meant to be.

docs/marmot-direct-messages.md claimed "An MLS room should never produce a
NIP-17 gift wrap for any reason". That premise is what made the guard look
complete, so it is corrected rather than merely amended, with the test to
apply when adding a kind to the exemption: can its recipient read a
kind:445? If so, it does not belong on this path.

**Not fixed, deliberately.** A refusal still leaves the payload unsealed
and head-blocking. That is now unreachable -- nothing else can queue a
payload against an MLS room -- but it remains a trap for whatever gets
refused next. Marking a payload refused needs a state the schema does not
have, so it is left for its own change rather than smuggled in here.

Verified: :composeApp:compileDebugKotlinAndroid succeeds. No test covers
this -- DatabaseChatRepository is Room-backed, and Room-backed code has no
unit test harness in this project.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-06 00:40:31 +02:00
Kgothatso Ngako
5abc37e463 docs: scope the jvm target, and separate it from testing the daos
Two questions arrived together -- whether Room's own testing guidance
applies to this project, and what desktop support would cost -- and they
turned out to have opposite answers. Both are now in docs/jvm-target.md,
phased, with the blocking work separated from the mechanical work.

**The expensive part is already done.** The four-deep native chain --
secp256k1 -> bitcoin-kmp -> lightning-kmp -> lightning-kmp-app -- already
builds for JVM, on every android build we do. The comment at
composeApp/build.gradle.kts:50 records the mechanism without drawing the
conclusion: lightning-kmp-core publishes no android variant, so our
android target resolves it to the *jvm* one, which pulls
secp256k1-kmp-jni-jvm desktop natives, which is exactly why the build has
to name the android artifact by hand. Read the other way round, every JVM
artifact in the chain is already compiled from source by the composite
build. A jvm target adds no cinterop, no C compilation and no new native
constraints. That was the part worth being afraid of, and it is finished.

**The blocker is one level down, and smaller than it looks.**
lightning-kmp-app/library declares 25 expects and implements them across
35 androidMain files. Its jvmMain holds exactly one: fibiprops.jvm.kt, the
Kotlin multiplatform library template's Fibonacci boilerplate, satisfying
two of the 25 -- both of them the template's own. So 23 actuals are
missing, which is why jvm() is commented out there
(library/build.gradle.kts:18), which is why it is commented out here
(composeApp/build.gradle.kts:46). Mantra cannot declare the target until
the fork does.

Six phases, ordered by that dependency. 0 build config; 1 the fourteen
mechanical phoenix actuals; 2 the three SQLDelight JDBC drivers and
NetworkMonitor; 3 key storage; 4 mantra's own sixteen expects; 5 the
desktop entry point. 1-3 are independent and parallelisable, 4 is where
the compiler finally checks the whole thing. Roughly a week to a
launchable build.

**Phase 3 has no day estimate, deliberately.** keyStoreEncryption /
keyStoreDecryption and their two graceful* wrappers delegate on android to
KeystoreHelper.kt -- 116 lines against AndroidKeyStore, StrongBox
attempted first and fallen back from, key material never leaving hardware.
Desktop JVM has no equivalent, so this is a decision rather than a port,
and the doc gives the three real options against what each actually
protects. A fixed-key JCEKS file is named there as a liability rather than
a stopgap: this is wallet seed material, and it lands on top of the
plaintext-key finding already open against this codebase. Recommended
sequencing is a passphrase-derived KEK with the desktop build marked
unsuitable for real funds, so phases 4 and 5 can proceed without the
security question being quietly treated as answered.

Two inherited mistakes are called out rather than carried forward. The old
Aux jvmMain put the database in java.io.tmpdir behind a TODO -- the doc
says not to inherit that in either phase that touches it. And
schedulePlatformLogic goes through WorkManager on android with no desktop
counterpart, so the doc asks for an explicit choice between a no-op and an
in-process coroutine, written down.

**The DAO answer is an appendix, because it is the opposite answer.** None
of the above is needed to test the DAOs, and burying that would have been
misleading. room3-runtime-android:3.0.1 already exposes the no-Context
inMemoryDatabaseBuilder(Function0<T>) overload, and
MantraDatabaseConstructor already supplies what it needs, so Room's
recommended host-machine form compiles in commonTest and runs under
testDebugUnitTest today. The one trap is native and is the secp256k1
problem mirrored: sqlite-bundled-android ships only android-ABI .so under
jni/, so a local unit test's JVM cannot load it and BundledSQLiteDriver
fails at construction; sqlite-bundled-jvm on the androidUnitTest classpath
is the fix. Robolectric neither helps nor is needed -- it cannot load
android .so on the host either.

Everything structural here was checked against the artifacts rather than
recalled: the Room builder overloads by javap on room3-runtime-android,
the two sqlite-bundled native layouts by unzipping both, and the
availability of room3-runtime-jvm, room3-testing, quartz-jvm and the two
SQLDelight drivers by request against the repositories this build actually
resolves from. The absence of android.* and java.* imports in commonMain,
and of any NFC reference from it, was likewise grepped rather than
assumed.

**Not verified: anything that requires compiling.** No jvm target was
turned on, nothing was built, and the day estimates are estimates. Phase 4
is where dependency-substitution surprises would surface if there are any,
and it is precisely the phase nothing here exercises.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-06 00:38:11 +02:00
Kgothatso Ngako
39eac61838 Merge branch 'mantra' into claude/long-running-chat-sync-8983dc
mantra had moved on ~30 commits, several of them in exactly this area — and it
turns out both branches independently found the same bug and drew the same
conclusion about the same filter.

**The overlap.** f38a5f1 fixed the three kind:1059 filters that named the wrong
pubkey, including the two `authors=[userPublicKey]` requests in NostrDao that
could never match a wrap signed by a throwaway key. This branch deleted those
same two blocks, inverting the same `if` to the `== null` case, for the same
reason. The code merged to the same shape; only the comments conflicted, and
they are combined.

**Nip17Filters wins, and the live subscription now defers to it.** ad3304a
extracted the inbox filter to one definition precisely because it had been wrong
in three call sites, with the no-`since` reasoning this branch arrived at
separately. Keeping a fourth copy inside LiveSubscriptionManager would recreate
the problem that commit exists to solve, so:

  - queueCatchUpSynchronization now calls Nip17Filters.inbox() instead of
    building an identical SynchronizationFilter with its own limit constant,
  - Nip17Filters gains liveInbox(), the same shape as a quartz Filter for a REQ
    rather than a SynchronizationFilter for the queue, and giftWrapFilter()
    defers to it.

Two types for one filter is not duplication worth removing — the queue stores
one and hashes it for computeId, a live subscription puts the other on the wire
— but they belong side by side, because drift here means one of them quietly
stops matching mail.

**ChatMessageListViewModel keeps this branch's resolution.** mantra had it
refresh our own inbox on open (Nip17Filters.inbox on our DM relays, purpose
"chat"); this branch removed that call entirely. Both were right when written,
and the merge is where the second becomes true: LiveSubscriptionManager holds
exactly that filter open on exactly those relays for the whole account and
reconciles it on every foreground, so opening a chat has nothing left to ask
for. The redundancy is now recorded in the comment where the branch used to be,
so it reads as superseded rather than dropped. Discovery — the kind-10050 lookup
for a participant we cannot yet address — is untouched, and the purpose is no
longer a conditional now that only one case reaches it.

The commonTest coroutines-test dependency arrived on both sides; the comment
gives both reasons.

Verified: 154 tests pass, both branches' suites included — Nip17FiltersTest and
the marmot direct-message suites alongside this branch's 46.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 23:58:40 +02:00
Kgothatso Ngako
bcdfd2ec94 Merge branch 'mantra' into claude/marmot-direct-message-type-7a0473
Twenty-two commits had landed on mantra since this branch left it, several
of them in the same files. Merged this way round so mantra stayed untouched
until the result compiled and its tests passed.

The migration had to be renumbered, and this is the conflict that mattered.
mantra is at database version 7 and already has its own 5.json -- for
MarmotInnerEvent.payloadEventId, nothing to do with direct messages. This
branch had also written a 5.json, for a different schema. Resolved by
restoring mantra's 5.json untouched and moving the direct message columns
to an AutoMigration(7, 8) with a regenerated 8.json. Taking either 5.json
over the other would have left every device validating a migration chain
against a schema it was never built from; keeping version = 5 would have
made a v7 install refuse to open at all.

The regenerated 8.json is two ADD COLUMNs and nothing else, same as before.

fromGroupEventResult was restructured on mantra: the kind switch moved into
applyInnerEvent, and a SubmissionEvent envelope now wraps nip30303 payloads.
Took that structure and re-applied the direct message branch ahead of it
rather than inside it -- a gift wrap is not a nip30303 payload to apply, and
what happens to it depends only on whether this device's key opens it, so it
does not belong in a function about applying submissions.

The isUserMessage fix was re-applied to the eight call sites mantra's
version has, up from the six it had here.

ChatMessageListViewModel and ChatRoomMessagingScreen took mantra's versions
with the composer state, the two renderings and the reply action layered
back on.

docs/README.md keeps both new rows and mantra's closing note about the
skipped-keys document.

108 tests pass, up from 50 here and 83 on mantra.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 23:45:08 +02:00
Kgothatso Ngako
c8cbd936f1 docs: record where the sync's safety net is, and where it is not
Two updates after the test pass.

long-running-sync.md gains a section naming what each test file pins and, more
usefully, the three things they cannot reach: NostrSocketClientImpl's reconnect
loop and ordered inbound (exercised only through their extracted arithmetic —
covering them wants a fake WebSocketSession), everything downstream of
saveNostrEvent (Room-backed, and there is no sqlite driver on the JVM test
classpath), and the app on a device. The manual checks stay the manual checks.

It also records that the tests were verified by mutation rather than by passing,
so the next person knows the assertions were confirmed to bite.

dead-code.md's line references are refreshed — the testability seams shifted
most of them — and it now says which commit they were correct at and to confirm
with the grep rather than trusting them. One entry added: the
DefaultNostrSocketClientFactory overload taking an explicit HttpClient has no
caller now that everything goes through the interface method.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 23:32:53 +02:00
Kgothatso Ngako
c24cbed390 docs: record what the coverage work found, and what it left uncovered
Three additions.

The decision the inbound path makes now has a name and a home --
MarmotDirectMessage.classify -- and the doc says why it is separate from
the filing of it: only the filing needs a database, so splitting them is
what lets the check that replaces MIP-03 be tested at all.

A security property found while writing those tests, which I had asserted
backwards. Relabelling a seal with another member's pubkey does not get as
far as the signature check: NIP-44 derives the conversation key from the
pubkey being claimed, so a relabelled seal is undecryptable by the person
it was encrypted for. The label is bound to the key rather than asserted
alongside it, and the outcome is a message the recipient genuinely cannot
read. verify() catches the narrower case of a seal altered after signing
in a way that survives decryption.

An honest list of what has no automated test and why -- the recipient
validation and the outbound id lookup (both need a database), the two
transcript renderings (no Compose UI test dependency in this project), and
anything touching a real MlsGroup. Better written down than rediscovered
by someone assuming a green suite means the path is covered.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 23:18:19 +02:00
Kgothatso Ngako
d110737f9a fix: keep a room's MlsGroup alive so a late message can still be read
Two events published in the same second reliably lose one of them. The
receiver stores the kind:445 and produces nothing from it -- no inner
event, no chat line, no error anybody sees, because MarmotGroupEvent is
written before the message is decrypted and so survives while everything
downstream silently does not.

Observed as a FROST signing session that never started on the receiver:
proposeSigning publishes the proposal and then the proposer's own nonce,
the relay handed them back in the other order, and the proposal was
dropped. The nonce is still sitting there filed against a session that
will never exist. The same bug ate a dialect earlier, which then took out
the artifact referencing it via a foreign key.

MLS is specified to tolerate this. RFC 9420 says a receiver that gets
generation N+1 before N keeps the intermediate keys so the older message
can still be read, and quartz's SecretTree does exactly that, in a
private skippedKeys map. What it does not do is persist it:
exportSenderStates() returns the ratchet positions only, so saveState()
drops the cache. NostrDao rebuilt the group from stored state for every
inbound event, so the cache was empty every single time, and generation N
arriving after N+1 failed `require(generation >= applicationGeneration)`
and was swallowed. Terminal -- the key is derived from a ratchet that has
moved past it, and nothing asks the sender to resend.

This keeps the instance alive instead. MlsGroupCache holds one MlsGroup
per room, and the inbound path goes through it, so skippedKeys survives
from one message to the next. That covers the case that actually bites --
a burst arriving in one sync, decrypted one after another against the
same tree -- which is what every bursty flow needs: proposeRitual sends
two, addArtifact sends two, and addChapter sends one per paragraph plus
one, of which only the ones arriving in ascending generation order
survived.

Reuse is conditional on the stored state still being exactly what the
cache last wrote. Sending a message advances the sender ratchet and saves;
so does adding a member. When that happens the cache rebuilds rather than
carrying on from a group that has been overtaken -- which is what keeps
this from being worse than no cache at all: the fallback is always the old
behaviour, never a diverged ratchet.

One lock per room, not one overall, because the group is mutable and
decryption advances it: two events for the same room decrypted at once
would corrupt the tree, and a busy room should not hold up a quiet one.

**This is a mitigation, not the fix.** It does not survive a restart, and
it does not survive another writer, so a long enough reorder still loses
the message. The fix belongs in quartz -- carry skippedKeys through
saveState/restore -- and quartz is a mavenCentral binary, not a fork, so
it cannot be made here. docs/mls-skipped-keys.md has the analysis, the
patch, the migration constraint on the persisted state format, and the
three ways to actually land it.

Not verified end to end: the proposal that exposed this cannot be
recovered, since its generation is already past, so confirming the fix
needs a fresh burst.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 23:09:53 +02:00
Kgothatso Ngako
635cef9311 docs: write down how a direct message travels, and what it costs
The reasoning behind this is not recoverable from the code, which is the
bar docs/README.md sets for having a document at all. Three things in
particular would otherwise have to be rediscovered by whoever changes this
next, and two of them are traps.

Why the wrap uses a throwaway key rather than the sender's own -- and what
that does not buy. It does not hide the sender from the group: MLS
authenticates every application message to a leaf, so the identity is
there regardless. What it costs is a carve-out in MIP-03's pubkey check
and the sender's ability to ever read their own messages back.

Why the check that carve-out removes is not a hole. The authorship claim
moves from the wrap's plaintext pubkey to the seal's verified signature,
bound to the MLS leaf that sent it -- strictly harder to forge than what
it replaced.

The one query that would broadcast one of these. What this builds is a
genuine, correctly signed NIP-59 gift wrap, indistinguishable from what
the NIP-17 path would be right to publish, and the only thing keeping it
off a relay is that it never becomes a GiftWrapPayload.

Written against what shipped rather than what was planned, so it records
two deviations. senderIdentity is resolved in NostrDao rather than added
to GroupEventResult.ApplicationMessage, because quartz is a binary
dependency here and the local checkout is a reference copy, not a build
input. And a failed validation drops the message and logs rather than
throwing, because the caller is inside storeNostrEvent's transaction.

The unbuilt parts are listed as absences rather than left implied: there
is no member picker, so a private message can only be a reply to one
somebody already sent, and nothing in the UI yet tells a user in words
that the group can see who they messaged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 19:10:22 +02:00
Kgothatso Ngako
c8c962e4f4 docs: inventory the unreferenced code in the sync and relay stack
Found while building the long-running sync. One item was orphaned by that
change; the rest was already dead and only became visible because the subsystem
was being read closely. Written down rather than deleted because several pieces
are one decision away from being wanted, and those decisions are not the sync
change's to make.

Every claim is "this identifier appears exactly once in composeApp/src, at its
own declaration", with the two things that method cannot see called out: Room
DAO methods are reached through generated code, and Compose entry points can be
invoked without a textual reference. The DAO cluster is flagged as the least
certain for exactly that reason.

Three findings are more than leftovers:

  - RelaysSocketManager.userRelays is a field nothing ever writes. The
    `userRelays` inside observeRelays is a different, shadowing local, so the
    single-argument publishEvent always takes its FALLBACK_RELAYS branch and the
    user's own relay list is never used for publishing. That is a bug wearing
    dead code's clothes, and the fix is to populate the field, not to delete it.

  - NostrPublisherRepository is entirely unreferenced, and it is the only
    consumer of CachingImportRepository.importEvents. RelayPool and
    RelaysSocketManager each take a cachingImportRepository parameter they store
    and never dereference, satisfied by NO_OP_CACHING_IMPORT_REPOSITORY — so the
    whole seam is a parameter passed from nowhere to nothing. Removing the
    publisher lets the interface and both parameters go with it.

  - sendAUTH is unused because NIP-42 is unimplemented, not because it is
    surplus. AuthMessage is parsed and dropped, so a relay answering CLOSED with
    auth-required is retried forever and can never succeed. Deleting sendAUTH
    means deciding against authenticated relays; that is worth doing on purpose
    or not at all. sendCOUNT and CountMessage are a similar matched pair — both
    go or neither, since a CountMessage cannot arrive if nothing sends a COUNT.

isRecommendedRelay on the two request entities is separated out as its own risk
class: never written, never read, but a Room column, so it wants a migration
rather than a delete.

Ends with an order to do it in, cheapest and least risky first.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 16:28:21 +02:00
Kgothatso Ngako
385c58ba7e docs: rewrite the sync note as what exists rather than what to build
The design landed across the six commits before this one, so the note is now
describing code. Reorganised around that: the reasoning that made it worth
writing is unchanged, but "the shape to build" is now "how it holds together"
and points at the classes, and the numbered traps have become properties of the
thing rather than warnings about a thing that did not exist yet.

Three sections earn their place after the fact:

  - the two timestamp decisions, which are the ones most likely to be "cleaned
    up" by someone who has not read this: no `since` on kind 1059 because our
    own wraps are stamped up to two days in the past, and no watermark on 445
    even though it would be safe, because `limit` already bounds the burst.
  - the four ways a group id can appear, which is why the group filter is
    derived from the room list rather than wired at the join sites.
  - "Not done", which was previously implicit in a staging plan: connectivity
    changes, NIP-42 AUTH, the collector-per-socket router the design originally
    called for, and the fact that the DM relay set is one relay.

The "suggested order" section is gone; git log is a better record of it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 16:20:08 +02:00
Kgothatso Ngako
178ddd0181 docs: write down how a long-running chat sync would work
Every chat sync today is a pull: a screen queues a request row, a pump drains
it, the relay answers, the subscription is closed. Nothing arrives between
pulls, so a message sent one second after EOSE waits for the next time someone
opens a screen.

This note works out what it takes to hold the two chat subscriptions open for
as long as the app is active — kind 1059 p-tagged to us, and kind 445 h-tagged
with every group we belong to — and, more usefully, what in the current
pipeline quietly assumes a subscription is short:

  - completeOnSubscriptionEnd finishes the flow at EOSE, which is what releases
    the slot and sends the CLOSE,
  - SUBSCRIPTION_TIMEOUT hard-kills anything still open at 120s,
  - subscriptionSlots is a Semaphore(4) shared with the backfill queue, so a
    permanent subscription is a permanently-held permit,
  - both saveNostrEvent overloads need a request row to attach provenance to
    and to flip to "processed",
  - and nothing in the app reconnects a dropped socket at all. That is
    invisible today only because every subscription is short and the next
    queued request re-opens the socket on its way out.

The design keeps the queue and its three pumps exactly as they are: live
subscriptions replace polling, not reconciliation. Negentropy stays the tool
for first login, the catch-up after a background gap, and "load older".

The group filter is derived from chatRepository.observeChatRoomListByPublicKey
rather than wired at each join site, because a group id can appear four ways
and only one of them (creating a group) is somewhere anyone would think to call
a subscribe function — being added arrives as a Welcome processed deep inside
NostrDao.storeNostrEvent. Observing the room list also closes the loop: a
Welcome lands on the live gift wrap subscription, a ChatRoom row is written,
the Flow re-emits, and the group filter widens without anyone opening a chat.

Two findings fell out of checking the details against our own code:

  - `since = now` on kind 1059 would silently drop messages. Gift wraps are
    stamped with TimeUtils.randomWithTwoDays(), so a wrap published now can
    carry a created_at two days in the past. Kind 445 uses TimeUtils.now() and
    can take a watermark — opposite treatment for the two kinds we care about.
  - the "sent-messages" filter (kinds=[1059], authors=[me]) cannot match
    anything, because gift wraps are signed with a fresh throwaway KeyPair().
    It is also unnecessary: createNip17ChatRoom puts the user in their own
    participant list, so we wrap a copy to ourselves and the account-wide
    #p=[me] subscription already picks it up.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 16:01:47 +02:00
Kgothatso Ngako
3dea07135c fix: add a group's whole membership in one commit, closing the epoch race
`MarmotOutboundDao.addMembersToChatRoom` stages every member with `proposeAdd` and
issues a single `commit()`. Both callers that know their membership up front now
use it: `SelectChatRoomTypeViewModel.inviteMembers` at room creation, and
`DkgRitualViewModel.inviteAdmins` for the #admins room.

Inviting one at a time created an epoch per member, and each of those commits
raced the previous member's welcome. MarmotInboundManager refuses future-epoch
messages outright, on both wire formats, with no queue and no replay -- so the
member who lost that race was silently stuck an epoch behind while the caller saw
a successful invite. Deriving isOneMemberInitialGroupCreation narrowed that window;
this removes it. No member ever has to process a commit for an epoch they were not
yet in, so there is no longer a race to lose.

One commit yields one welcome: `buildWelcome` emits an EncryptedGroupSecrets per
added member and each joiner finds its own entry by key package reference. The blob
is shared, delivery stays per peer, because each welcome event is tagged with that
peer's key package.

## Why this needed no schema change

Batching at creation time means the single commit happens while the group is still
only its creator, which takes the immediate-welcome branch: nothing is broadcast
and MarmotCommitResult is never written. The bookkeeping that assumes one peer per
commit is simply not on this path.

So the batch is taken only when `members().size == 1`, and anything else falls back
to inviting sequentially -- correct, if not ideal. Batching into an established
group would take the deferred branch, where `peerKeyPackageEventId` is singular and
the ack-triggered delivery in DatabaseNostrRepository expects one welcome; making
that work needs a list there and a fan-out on acknowledgement. Nothing currently
adds several members to an established group, so that is left outstanding and
documented rather than speculatively built.

The group state is persisted after `commit()` and before any welcome goes out, so a
crash between them leaves the group at the epoch the welcomes describe rather than
one behind it.

## Reporting

Members with no published key package still cannot be added -- a Marmot invite
needs one -- and are now returned alongside any that failed to receive their
welcome, rather than the two being conflated. Both still only reach the log; the
coordinator is not yet told.

docs/marmot-membership.md is updated in the same change: batching moves from
outstanding work to described behaviour, with the schema constraint that shapes it
and the remaining fan-out work recorded. The note about sequential invites is
narrowed to where they still happen.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 15:02:06 +02:00
Kgothatso Ngako
8fc1c9e650 fix: stop deferring the first invitee's welcome behind a commit nobody needs
`inviteMember` now works out for itself whether the group it is adding to has
anybody to inform:

    val isOneMemberInitialGroupCreation = mlsGroup.members().size == 1

read before `addMember` advances the tree. The parameter is gone from the
signature and no caller passes it any more.

Callers were the wrong place for this decision and both of them got it wrong.
`inviteMemberToChatRoom` hardcoded `false`, and `ChatRepository.inviteMember` did
not expose it at all, so every invite made through a group -- room creation in
SelectChatRoomTypeViewModel, and the #admins room -- took the deferred-welcome
path. That includes the first invite, when the group is still only its creator, at
which point:

  - the commit has no audience. No other member exists, and nobody outside the
    group can decrypt it, so it is noise on the relay.
  - the welcome is then withheld until a relay acknowledges that noise. If the ack
    never lands, the first invitee receives nothing at all.

Only createMlsDirectMessageChatRoom passed `true`, and only because a DM has
exactly one invite. A group of n has one such invite too -- the first -- and it was
not getting it.

The condition is right at any size, not just for DMs: "the group has nobody to
inform" is true exactly once. Invite two sees one member who must advance, invite
three sees two, and so on. Their commits are encrypted with
`commitResult.preCommitExporterSecret`, the epoch the earlier invitees received in
their own welcome, so they can decrypt and advance. `members()` skips empty leaves,
so this also stays correct for a group that has had members removed.

## What this does and does not fix

It removes a pointless commit and, with DefaultDMRelayList now a single relay, a
single point of failure sitting in front of every group's first member.

It also narrows a silent race rather than closing it. MarmotInboundManager refuses
future-epoch messages outright on both wire formats -- no queue, no replay -- so a
commit arriving before its recipient's welcome is dropped and that member never
advances, while the coordinator sees a successful invite. Previously both commits
went out before either welcome; now welcome 1 is sent before commit 2 exists, so
the first invitee is already at the right epoch. For n >= 3 the window between
welcome 1 and commit 2 remains.

Closing it needs the adds batched into one commit, which is the outstanding work
described in docs/marmot-membership.md. That doc is updated here to describe the
derived flag as current behaviour rather than a proposal, and to keep batching as
the remaining item.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 14:44:28 +02:00
Kgothatso Ngako
b99cb8fcd5 docs: write down the shared-key subsystem and how Marmot membership fails
First docs in the repo -- README.md is still the stock KMP template. Three
documents plus an index, covering the parts whose behaviour is not recoverable by
reading the code: where the reasoning lives in a protocol, where a failure mode is
silent, or where a decision looked arbitrary and was not.

marmot-membership.md is the one that earns its place. Everything about adding a
member compiles, the invite reports success, and a member simply never appears --
and the reason is never in the invite code. It records that
inviteMemberToChatRoom hardcodes isOneMemberInitialGroupCreation = false and that
ChatRepository does not expose it, so every group invite takes the deferred-welcome
path including the first, when the group is still just its creator and the commit
has no audience at all. Then why that is silent rather than noisy:
MarmotInboundManager refuses future-epoch messages outright, on both wire formats,
with no queue and no replay, so a commit arriving before its recipient's welcome
is dropped and that member never advances. EPOCH_RETENTION_WINDOW retains past
epochs and does nothing for messages from ahead. Three options are set out with the
per-invite correctness table, including the honest limit that the recommended one
narrows the race without closing it.

shared-key-derivation.md argues why the paths are not BIP32 -- no chain code
exists, hardened derivation is impossible rather than unimplemented, and a FROST
tweak takes the scalar as input so the chain code leaves the problem entirely. It
records the x-only serialisation trap avoided by choosing the scalar directly, and
states the rule that must not be broken: never reconstruct a derived key in the
clear, because k = k' - t hands over the group key rather than one derived key.

shared-key-ceremony.md covers the seven kinds, the three approval gates and why
the coordinator's aggregations are deliberately not among them, faults as values
rather than exceptions, and the transcript's idempotency-by-construction. It also
writes down the invariant that produces no error when broken: pendingApproval must
mirror the gates in advance, or the screen offers an approval that does nothing --
or none while the ritual sits still.

Every factual claim was checked against the source rather than recalled, which
turned up one correction worth having: there are two future-epoch refusals, for
PrivateMessage and for Commit, so the drop covers both wire formats and not just
one.

Each document leads with the failure mode rather than the architecture, on the
grounds that a failure is what sends somebody to docs in the first place, and each
lists its known gaps -- including that none of this has run on a physical device.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 14:39:19 +02:00