Commit Graph

33 Commits

Author SHA1 Message Date
Kgothatso Ngako
dff41d417d feat(frost): move a signing session's per-event columns onto FrostSigningItem
Phase 1 of docs/frost-batch-signing.md, which is added here as the plan the
next phases follow. Schema only: a session still signs exactly one event, the
wire is byte-identical, and every existing test passes on the moved columns.

## What moved, and why it had to

A batch of k events is k independent FROST instances sharing a signer set, not
one signature over k messages. That is forced rather than chosen: a Schnorr
partial signature is `s = k + e·x` with `e = H(R‖P‖m)`, so two messages under
one nonce R give two equations in one unknown and the secret share falls out.

So the five columns that enter that equation -- unsignedEventJson, eventId,
nonceRandom, aggregatedNonce, signature -- move to a child table keyed
(sessionId, itemIndex). What stays on FrostSigningSession is everything outside
it: the ceremony, the threshold, the derivation path, the signer set, and the
one approval.

itemIndex is protocol rather than presentation -- nonces and partial signatures
are joined positionally against it -- so getItems() orders by it and nothing
re-sorts. Spelled itemIndex rather than index to keep hand-written queries free
of backticks.

No itemCount column. The count is a COUNT(*), for the same reason signerIds is
derived from the ceremony's participant order rather than stored: a
denormalised count is one more thing that can disagree with the rows.

## Migration 9 -> 10

Manual, not auto: Room can create the table and drop the columns but cannot
copy between them, and the copy is the whole point. A session in flight at
upgrade holds its nonce seed and the aggregate it is already signing against,
and neither can be regenerated -- losing either makes the next pass derive a
different nonce for the same message and publish a second partial signature
over it, which is the extraction case. Both are copied verbatim into item 0, so
an in-flight session resumes as though nothing happened.

Removing the columns uses ALTER TABLE DROP COLUMN rather than the usual
create-copy-drop-rename rebuild. FrostSignerMessage and FrostSigningItem both
reference FrostSigningSession(id) ON DELETE CASCADE, and DROP TABLE fires
cascades -- with foreign keys enforced the rebuild would delete every signer
message and every item just written. Whether it does depends on Room disabling
foreign keys around migrations, which is not worth depending on when
DROP COLUMN cannot go wrong. It needs SQLite 3.35 and unindexed,
unconstrained columns; these five qualify, and getRoomDatabase pins
BundledSQLiteDriver on every platform.

## Invariants established here for the phases that follow

- signerIds and every item's aggregatedNonce are one write-once unit, applied
  by applyAggregate() -- items first in one transaction, then the session, so
  "some items aggregated" is unreachable and signerIds != null stays the gate.
- Signatures likewise, via applySignatures(); isSigned() counts rows instead of
  reading a flag.
- complete() verifies every signature before applying any event, so a batch is
  all-or-nothing rather than half-filed.
- itemsOver() gives each item its own 32 bytes of seed. Independent seeds mean
  an off-by-one in index handling produces a session that fails to aggregate
  rather than one that signs two messages under a single nonce.

signedEvent() and isAwaitingApproval() now take the item(s) rather than the
session, which propagates to the repository, the view model and the screen.
advance() reads items.first() and Phase 2 turns that into a loop.

## Tests

- FrostSigningSessionDaoJvmTest: index ordering, single-item read, upsert
  replacing rather than accumulating, signed-item counting, cascade delete.
- FrostSigningItemMigrationJvmTest (new): the backfill against a real v9
  database, asserting the seed and aggregate values survive -- not merely that
  a row appeared -- plus the exact column lists Room will check at open time.
- 338 jvmTest and 217 testDebugUnitTest pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-06 04:33:46 +02:00
Kgothatso Ngako
a8f6638325 feat: sign a room's key state into being, at the room's own key
Two changes that turned out to be one. A room's key state stops being something
its creator announces and becomes something the group signs, and every FROST
signature moves from the group's root threshold key to the key derived at the
room's own path -- which is the room's id. The second is what makes the first
worth having: a key state is now signed by the very key it names.

Supersedes the announcement introduced in a909108, and changes the author of
every event the group signs, including the artifacts of 786c060.

## The key state is proposed, not announced

a909108 had the room's creator write the GroupKeyState row, say so in the room
on kind 30326, and every receiver keep it if the room's id rederived from the
key it named. That check was sound and is still here -- a state that does not
rederive its own room is dropped, whoever sent it -- but it left the first thing
a group ever does as the one thing a single member decides alone.

So the key state goes through the door everything else the group says goes
through. GroupKeyStateManager.announce becomes propose, which opens a
FrostSigningEvents.PROPOSAL over an unsigned 30326 and writes no row. The state
comes into existence when a quorum has signed it, on every device at once,
applied by FrostSigningManager.complete like any other signed proposal:

    creator  --[ 30320 proposal over an unsigned 30326 ]-> everyone
    ...members approve, nonces, signer set, partials, aggregate...
    everyone --applies the signed 30326 locally-->  GroupKeyState row

Nothing waits on it. Between creating a room and that session completing there
is no state to read, and completedKey's rederivation scan -- kept from before
the table existed -- is what keeps the room signable in the meantime, including
for the key-state session's own members. That is the only reason a bootstrap
here does not deadlock, and the scan's doc now says so rather than describing
itself as legacy.

proposeSigning gains an optional `key`, for the one caller that cannot be asked
which key the room signs with because establishing that is its whole job. It is
honoured only if this device actually holds a share of it, so naming a ceremony
cannot talk a session into signing with material it has not got.

## Everything signs as the room, not as the group's root key

unsignedEventOf and advance now build from SharedKeyDerivation.derive at the
room's path instead of TweakCache.create on the bare threshold key. Both halves
had to move together: a signature aggregates against whatever the cache carries,
so the cache and the author on the event have to be the same derivation or
nothing verifies.

Since marmotGroupId(K, path) *is* derive(K, path).hex, the pubkey on every event
a room signs -- dialect, artifact, chapter, key state -- is now that room's id.
A reader checking one needs no lookup at all: the key they expect is the id of
the room they found it in. marmotGroupId's doc now carries that second meaning,
and there is deliberately no second name for the value; "the room's id" and "the
key it signs as" are one function because they are one key.

The path is resolved by FrostSigningManager.signingPath and never taken from a
proposal, because it decides which key the group signs as -- a proposer able to
choose it could have every signer put their share behind an author of the
proposer's choosing. Three candidates in descending order of knowledge (the
room's GroupKeyState, the path in its MIP-01 description, the app's default),
and one is accepted only if walking it reaches the room's id, which makes the
resolution self-checking rather than trusting. acceptProposal runs the same
resolution independently on every device.

Null is a real answer, not a failure: completedKey will still find a key for a
room that was never derived from it -- a ceremony held in that very room, the
fallback kept for rooms the app no longer makes -- and such a room has no key of
its own to sign as, so it signs as the threshold key, which is what it always
did.

## What a receiver now checks

GroupKeyStateManager.stateFrom asks two independent questions, and a state has
to answer both:

  - Is it true? The room's id is the key derived at the path, so a state that
    does not rederive its own room names a key the room was not made from.
    Unchanged, and still the half that safety rests on. It knows nothing about
    who is speaking, and that is deliberate: a member with no share can state a
    true state and it is still true.

  - Did the group say it? GroupKeyStateEvent.isSignedByGroup: the author must be
    the key the content walks to at the path in the tags, the id must hash the
    fields sitting next to it, and the signature must verify. Since that walk is
    the room's id, a passing state is signed by the room it is about.

The second does not make a state truer -- the derivation already settled truth.
It makes the record of what a room signs with a thing a quorum agreed to. The
practical effect is that a true state nobody signed is now refused, which is the
behaviour change worth knowing about: an unsigned 30326 from an older client is
stored as an inner event and dropped as a state.

## Restart safety, and a nullable column

FrostSigningSession gains derivationPath, and the database goes to v9 on an
auto-migration. It is an input and is stored for the same reason nonceRandom is:
the cache is rebuilt on every pass of advance, and a session that resolved a
different path after a restart would regenerate a different nonce from the same
seed -- publishing a partial signature against an aggregate nobody else
computed.

Nullable, meaning no derivation at all: the untweaked threshold key. That is
both the honest answer for a room not derived from the key and what sessions
predating the column read back as, so a session caught mid-flight by the
migration finishes under the key it began under rather than switching between
two of its own rounds.

SharedKeyDerivation.derive now takes its key back out of the cache rather than
from the point, so an empty path is a real answer equal to what a session
created from that cache signs against. No behaviour changes for a non-empty
path, where the walk overwrites it on the first step.

## One place that files a key state

ChatMessage.applyInnerEvent records it, which it must: the signed 30326 reaches
every device through applySignedEvent, and the branch there previously returned
null and dropped it. The now-duplicate dispatch in NostrDao is removed, so the
locally applied signature and any wire-borne 30326 take the identical path.
Still no chat line -- standing state, and the session already wrote the
transcript of it happening.

## Elsewhere

DkgRepository.announceGroupKeyState becomes proposeGroupKeyState, taking the
room and returning the signing session rather than the state, since the state is
not what the call produces any more. DkgRitualViewModel calls it after members
are added, unchanged and for the unchanged reason: adding them commits a new
epoch, and a proposal published before it reaches nobody who could sign it.

FrostSigningScreen describes a key-state proposal as the group's shared key with
its path and ceremony, rather than "Event of kind 30326" -- a member deciding
whether to sign should be shown the thing.

## Tests: 217, 0 failures

  - GroupKeyStateTest is rewritten around real quorum signatures from
    Frost.trustedDealerKeygen. New: a true state nobody signed is dropped, a
    member's own signature over one is dropped, one group signing about another
    group's key is dropped, a state edited after signing is dropped, a state
    signed at the wrong path is dropped, and the room signs as its own id.
  - SignedArtifactTest pins that an artifact's author is the room it was signed
    in, and explicitly not the group's root key.
  - FrostSigningRoundTest runs both rounds against the tweaked cache now, which
    is the part most likely to be silently miswired -- a badly built cache
    produces a signature that simply fails to verify, on every device, quietly.
  - SharedKeyDerivationTest pins the migration contract: a walk of no steps
    lands on the threshold key, and a null derivationPath reads back as that
    empty walk rather than as the default path.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-06 03:18:58 +02:00
Kgothatso Ngako
607ef72bc3 Merge branch 'mantra' into claude/marmot-group-reindex-events-96a0d0
# Conflicts:
#	composeApp/src/commonMain/kotlin/press/mantra/compose/database/dao/NostrDao.kt
#	composeApp/src/commonMain/kotlin/press/mantra/compose/database/model/ChatMessage.kt
2026-09-06 01:52:10 +02:00
Kgothatso Ngako
925099125b feat: read a room's group events again when they arrived out of order
Relays impose no ordering, so a kind:445 can turn up before the group can
read it: an application message encrypted under an epoch whose commit has
not landed, or a commit for an epoch ahead of the local one. Both are
stored and then dropped -- MarmotInboundManager refuses an out-of-epoch
commit precisely so it does not half-mutate the group -- and nothing goes
back for them once the missing event fills the gap. The message is on
disk, readable, and never read. A "Reindex Events" button at the bottom
of the group's detail screen is that second look.

Only events with nothing to show for them are replayed: no chat line at
all, or one of the two placeholder types. A room where nothing went wrong
is left exactly as it was, which is what makes the button safe to press
on a hunch. Passes repeat while a pass recovers something, because
created_at order is not epoch order and a commit recovered by one pass is
what lets the next read the messages that were waiting on it.

**Replaying was not safe as it stood.** Every row the path writes is keyed
on an event id and upserts in place -- MarmotGroupEvent, MarmotInnerEvent,
and the nip30303 entities -- with one exception. ChatMessage's primary key
is autogenerated, so writing a freshly built line always inserts, and a
re-read would have left the room showing each recovered message twice,
once as "Undecryptable Message" and once as itself.
ChatMessage.reconcileMarmotLine matches on the group event id instead, so
a re-read is an update, and refuses to let a placeholder overwrite a line
that says something. That last rule is what protects the line this device
wrote on the way out for a message it sent: our own kind:445 cannot be
read back, since the sender ratchet has consumed the generation, and
without the rule a replay would have replaced our words with
"Undecryptable Message".

The MLS group itself was already safe to replay against, which is worth
saying because it is the part that looks dangerous: a commit behind the
current epoch is rejected as a duplicate before it touches the group, one
ahead is refused, and a consumed ratchet generation throws before
mutating anything. The exception was quartz's EpochCommitTracker, which
does not dedupe and only empties when a commit applies -- so replaying a
held commit just grew the list and left it pending forever.
forgetPendingCommits drops the room's entries first, and the sweep feeds
the events back in the order CommitOrdering picks a winner in, so a
contested epoch resolves the same way it would have on every other
device.

**What is testable, and what is not.** The DAO is not: testDebugUnitTest
is plain JVM and Room's in-memory builder wants an Android Context. So
the two pieces carrying decisions are lifted out where they can be run
without one -- MarmotReindexSweep for the stopping rule, and
reconcileMarmotLine for which of two lines wins -- and the DAO is left as
query, sweep, write. The filter tests pin why the query's `tags LIKE` is a
prefilter and not a test: an event belonging to another room can mention
this one in a q tag, and its own h tag is what rejects it.

**Not recovered by any of this.** A message whose key is gone -- one the
ratchet has already advanced past, or one from an epoch predating this
device's join. And events that never reached disk at all: storeNostrEvent
is a single transaction, so a kind:445 arriving before its room exists
rolls back its own insert along with the failed indexing, and only a
re-sync brings it back.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-06 01:49:57 +02:00
Kgothatso Ngako
6aff34c5c7 Merge branch 'mantra' into claude/artifact-frost-signing-proposal-a414f6 2026-09-06 01:38:08 +02:00
Kgothatso Ngako
786c0602da feat: sign an artifact into the library instead of submitting one
Adding an artifact no longer creates one. It opens a signing session over
an ArtifactEvent, and the artifact appears -- on every member's device at
once, authored by the group's shared key rather than by whoever typed it
-- when enough members have signed. The same trade the dialects made: a
submission says "I am putting this in front of the group" and the group's
only recourse afterwards is social, while a signature is the group saying
it and it takes a quorum to say. A library is the group's.

**The first version.** This is the part the dialect had no answer for. An
artifact was creating an initial ArtifactVersion as a second submitted
event, and that cannot survive the change: a chapter attaches to a
version rather than to an artifact, so an artifact without one is inert,
but a version cannot be submitted before the artifact it points at
exists, cannot have its own quorum without costing a second signing
session per form, and cannot be invented locally -- an invented id
differs on every device, so members would silently disagree about which
version a chapter hangs off while every screen showed the same artifact.

So the label rides on the artifact as an `artifactVersion` tag and the
row is derived from the signed artifact's own fields when it is applied.
Same bytes in, same row out, everywhere. It is a rumor, because nobody
signed it; what the group signed is the artifact that declares it.

**What went away.** MantraDao.addArtifact and its way up through the
repository. Nothing called it once the screen proposed instead, and
leaving a path that authors an artifact under a member's key while the UI
insists on a quorum would have double-created the version besides.

**Tests.** Three files, and each was checked against a broken
implementation rather than only against a working one: deriving the
version from the clock, dropping the label from the proposal, authoring
the derived row as its reader, and losing the signature on the way out of
the session are all caught. SignedArtifactTest runs a real 2-of-3 quorum
over an actual proposal, because the claim worth holding -- the row is
the group's, and carries proof of it -- is invisible when it breaks.

Not covered: applyInnerEvent's two upserts, which need a database no test
here stands up, and AddArtifactViewModel, which is plumbing across two
dispatchers over a template the tests already pin.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-06 01:37:58 +02:00
Kgothatso Ngako
024da99404 test: pin what a member who never took part needs to finish a session
e22a8ae hoisted completion above the approval gate on the strength of one
claim: closing a session needs nothing secret, and nothing the member would
have had to publish. Were that false -- were the aggregated nonce, the
signer set or a share needed to check the result -- the gate would have to
stay where it was, and the member a quorum did not need would go on being
asked to sign something already signed. Nothing checked the claim.

FrostSigningCompletionTest builds a real 2-of-3 signature from members 0
and 1, then works entirely from member 2's row: never approved, not in the
signer set, aggregatedNonce and signerIds deliberately null. From that
alone it pins that they can verify what the group signed, that the finished
event is the proposed one unaltered rather than rebuilt or rehashed, that
there is no finished event before the signature arrives, that the arrived
signature is what stops the session asking, and that a signature over a
different event is refused -- which is what the check in complete() is for.

Both new assertions about isAwaitingApproval were mutation-checked: with
the `signature != null` guard removed, exactly two tests fail and the rest
of the suite still passes, so they guard the change rather than restating
it.

Not covered, and not coverable here: advance() itself -- that the branch
fires on an inbound SIGNATURE rather than stopping at the gate. It is
Room-backed, and this project has no harness for that (no Robolectric, and
the in-memory builder's android actual needs a Context). The pure half of
the claim is what this pins instead.

Also corrects e22a8ae's message, which said sixteen new tests. It was
eleven: eight in TranscriptRequestStateTest and three added to
FrostSigningSessionTest, which has eight in total.

Verified: :composeApp:compileDebugKotlinAndroid succeeds, and
:composeApp:testDebugUnitTest passes -- 176 tests across 24 classes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-06 01:26:52 +02:00
Kgothatso Ngako
b87e6e4ed5 Merge branch 'mantra' into claude/frost-proposal-review-visibility-8f6fab 2026-09-06 01:13:47 +02:00
Kgothatso Ngako
e22a8ae4cd fix: stop asking a member to review a signature the group has settled
The transcript's "Review" affordance is a promise: tapping it leads to a
decision still there to be made. For a FROST signing proposal it was only
ever withdrawn one way -- and a proposal can be processed three.

**How a request was closed.** RitualNotice drops the tint and the call to
action when the request is answered, and a request counts as answered when
the step it asked for has since been published by this device:

    FROST_REQUEST_FULFILMENTS = mapOf(TYPE_FROST_APPROVAL_NEEDED to TYPE_FROST_NONCE)

Approving publishes a nonce, so approving closes it. Nothing else does.

**Declining.** decline() fails the session and broadcasts a FAILURE. It
publishes nothing of the member's own, by design -- a refusal is a refusal.
So no fulfilment line is ever written, and the request went on asking, in
primary tint, for a decision the member had already made. Tapping it
reached a screen with no buttons on it, which was the screen being right.

**A quorum that did not need them.** A t-of-n key finishes without
everybody. The coordinator takes the first t nonces, and a member whose
phone was in a pocket is simply not among them -- but advance() returned at
the approval gate on their device, so the arriving SIGNATURE was stored and
nothing was done with it. Their session sat at COLLECTING_NONCES forever.
The request stayed lit, the screen still offered Sign and Don't sign, and
both answers were wrong: a nonce nobody was waiting for, or a refusal that
would flip a COMPLETE session to FAILED on every device and announce
"Nothing was signed" to a group holding the signature. fail() writes the
stage with update() rather than moveTo(), so that last one was reachable.

**The transcript.** A request is now closed by being *answered* or by being
*settled* -- a frostComplete or frostFailed line after it. The two are kept
apart deliberately. Answered keeps the tick; settled does not, because the
member never answered and crediting them with a signature they refused, or
were never asked for, is worse than the summons was. Both rules moved out
of the composable onto ChatMessage, where they are stated once and tested.
Settlement is signing-only: a ceremony step can only be taken or waited
for, so a DKG request has no equivalent and reading one from a signing
session's end would drop a summons the ritual is still stalled on.

**The session.** The transcript alone could not close the third case: the
device that never approved wrote no terminal line to read. advance() now
completes on a signature that has already arrived, ahead of the approval
gate rather than below it. That gate is there to keep this device's own
material off the wire, and finishing puts none there -- it verifies the
aggregate, applies the event and announces, all from what is already
stored. Everything it now skips on that path is work the signature made
pointless anyway: a late nonce, a partial signature nobody will aggregate.

Three things follow. isAwaitingApproval reports false, so FrostSigningScreen
hides the buttons -- it now asks the manager rather than re-deriving the
rule, which had drifted into a second copy of it. A late "Don't sign"
cannot abandon a signature that exists. And the signed event finally lands
locally for a member who never approved: applySignedEvent sat below the
gate and was being skipped, so a dialect the group signed without them
never reached their store.

Verified: :composeApp:compileDebugKotlinAndroid succeeds, and
:composeApp:testDebugUnitTest passes -- 165 tests, 16 of them new. Eight
cover the transcript rules against a hand-built row list; eight cover
isAwaitingApproval, including the settled-signature case. What stays
uncovered is advance() itself, which is Room-backed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-06 01:13:42 +02:00
Kgothatso Ngako
02117643c4 fix: send a group event because it was queued, not because the chat mentions it
No FROST signing message has ever reached another participant. The proposal
was built, MLS-encrypted, wrapped under the exporter secret, signed as a
kind:445, written to NostrEvent and MarmotGroupEvent, and its queue row
marked processed -- and then never handed to a relay, by a branch that was
never about delivery at all.

**The gate.** The tail of MarmotOutboundDao.encryptAndSendMarmotInnerEvent
looked up the transcript row for the queued rumor and did everything else
inside it:

    val chatMessageOrNull = database.chatMessageDao()
        .getChatMessagesByMarmotInnerEventId(marmotInnerEvent.id)
    chatMessageOrNull?.let { chatMessage ->
        ... relation, marmotGroupEventId ...
        val ids = database.broadcastNostrEventRequestDao().insert(...)
    }

The BroadcastNostrEventRequest rows are the only thing that puts a kind:445
on a relay -- observeBroadcastNostrEventRequestsByStatus("pending") is what
the broadcaster watches, and nothing else inserts them for this path. So the
question "does the chat have a line for this?" was silently answering the
question "should the group receive this?".

**Why FROST always lost.** A signing message has no ChatMessage by design.
FrostSigningManager.broadcast queues the rumor alone, and announce() writes
its milestone lines with marmotInnerEventId = null on purpose: each device
writes its own transcript from the messages it has already received, so the
lines cost no traffic and cannot disagree with the session they describe.
The inbound half states the same intent from the other side --
ChatMessage.applyInnerEvent returns null for every FrostSigningEvents kind,
because a row there would be a second, worse account of what the manager
already narrates.

That is every kind in the family, not just the proposal: nonces, the signer
set, partial signatures, the finished signature and the failure notice all
go through the same broadcast(). A session could not have completed even if
a proposal had somehow arrived.

**GroupKeyStateManager.announce had it too.** Same shape, same silence: a
room's kind:30326 announcement of which key it signs with was queued,
encrypted and dropped. a909108 added it so members would stop rederiving;
no member has ever received one.

**Why the neighbours worked, and hid it.** The DKG rides NIP-17 gift wraps
through a different path entirely, so a room could finish a ceremony, hold a
real shared key, and report canSign() == true with the signing transport
dead beneath it. nip30303 submissions work because MantraDao.sendMarmotInnerEvent
pairs every queued rumor with a ChatMessage carrying its id -- not as a
delivery mechanism, just because a submission is also something a member did.
FROST was the first traffic to use the group path without a chat line, which
is why this reads as a FROST bug and is not one.

**Why it went unnoticed.** Nothing failed. The coordinator's own device is
fully convinced: proposeSigning writes the session, announceStarted puts a
line in the chat, advance() runs, and publishOwn records the coordinator's
own nonce and announces that step too. From the proposer's side a session
nobody else can see is indistinguishable from one waiting on slow peers.

Unlike 65e4a3a, the queue did not block. marmotGroupEventId is set before
the transcript lookup, so the row left the queue cleanly and the next one
was picked up. Every message was lost individually, in silence, with no
backlog to notice.

**The fix.** The broadcast insert is hoisted out of the branch, and the
decision it was tangled with is lifted into MarmotDelivery.plan: given a
group event, a relay list, and a chat message or null, what has to be
written. A group event is sent because it was queued; a chat line is linked
because a member said something. The DAO now computes that plan and executes
it, with the insert as a plain unconditional statement ahead of the
bookkeeping that legitimately does depend on there being a line.

The extraction is not decoration. encryptAndSendMarmotInnerEvent is
Room-backed and cannot be stood up in a unit test, which is exactly how the
gate survived; separating the decision from the filing of it is the same
move MarmotDirectMessage.classify exists for, and for the same stated
reason.

**Tests.** MarmotDeliveryTest, six of them. The two that matter are "a
signing proposal goes out, though nothing in the chat points at it" and "the
send does not depend on the transcript", which asserts the broadcast list is
identical with and without a chat message. The rest pin the supporting
facts: one request per relay naming the event, every request written pending
because that is the only status the broadcaster looks at, the linkage that
does depend on a chat line, and an empty relay list as the sole legitimate
way to produce an empty broadcast list -- so that an empty list always reads
as "nowhere to send it" and never as "nothing to send".

**Not covered, deliberately.** These pin the decision, not the call site. Re-
nesting the insert inside chatMessageOrNull?.let would leave MarmotDelivery
correct and every test passing. Closing that needs the DAO itself under
test: BundledSQLiteDriver is on the classpath and getInMemoryDatabaseBuilder
exists, but its android actual wants a real Context, testDebugUnitTest is
plain JVM, and there is no androidUnitTest source set or Robolectric. That
is its own change, not one to smuggle in here.

Verified: :composeApp:compileDebugKotlinAndroid succeeds; 160 tests pass,
154 before these six. The inbound half was read rather than assumed --
NostrDao dispatches FrostSigningEvents kinds to processSigningPayload,
inbound rumors are stored with marmotGroupEventId set so they cannot re-enter
the outbound queue, and the out-of-order replay path is intact. Outbound was
the only break. That two participants now actually see a proposal is
inference from the code, not an observation: it wants two devices.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-06 01:06:05 +02:00
Kgothatso Ngako
39eac61838 Merge branch 'mantra' into claude/long-running-chat-sync-8983dc
mantra had moved on ~30 commits, several of them in exactly this area — and it
turns out both branches independently found the same bug and drew the same
conclusion about the same filter.

**The overlap.** f38a5f1 fixed the three kind:1059 filters that named the wrong
pubkey, including the two `authors=[userPublicKey]` requests in NostrDao that
could never match a wrap signed by a throwaway key. This branch deleted those
same two blocks, inverting the same `if` to the `== null` case, for the same
reason. The code merged to the same shape; only the comments conflicted, and
they are combined.

**Nip17Filters wins, and the live subscription now defers to it.** ad3304a
extracted the inbox filter to one definition precisely because it had been wrong
in three call sites, with the no-`since` reasoning this branch arrived at
separately. Keeping a fourth copy inside LiveSubscriptionManager would recreate
the problem that commit exists to solve, so:

  - queueCatchUpSynchronization now calls Nip17Filters.inbox() instead of
    building an identical SynchronizationFilter with its own limit constant,
  - Nip17Filters gains liveInbox(), the same shape as a quartz Filter for a REQ
    rather than a SynchronizationFilter for the queue, and giftWrapFilter()
    defers to it.

Two types for one filter is not duplication worth removing — the queue stores
one and hashes it for computeId, a live subscription puts the other on the wire
— but they belong side by side, because drift here means one of them quietly
stops matching mail.

**ChatMessageListViewModel keeps this branch's resolution.** mantra had it
refresh our own inbox on open (Nip17Filters.inbox on our DM relays, purpose
"chat"); this branch removed that call entirely. Both were right when written,
and the merge is where the second becomes true: LiveSubscriptionManager holds
exactly that filter open on exactly those relays for the whole account and
reconciles it on every foreground, so opening a chat has nothing left to ask
for. The redundancy is now recorded in the comment where the branch used to be,
so it reads as superseded rather than dropped. Discovery — the kind-10050 lookup
for a participant we cannot yet address — is untouched, and the purpose is no
longer a conditional now that only one case reaches it.

The commonTest coroutines-test dependency arrived on both sides; the comment
gives both reasons.

Verified: 154 tests pass, both branches' suites included — Nip17FiltersTest and
the marmot direct-message suites alongside this branch's 46.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 23:58:40 +02:00
Kgothatso Ngako
bcdfd2ec94 Merge branch 'mantra' into claude/marmot-direct-message-type-7a0473
Twenty-two commits had landed on mantra since this branch left it, several
of them in the same files. Merged this way round so mantra stayed untouched
until the result compiled and its tests passed.

The migration had to be renumbered, and this is the conflict that mattered.
mantra is at database version 7 and already has its own 5.json -- for
MarmotInnerEvent.payloadEventId, nothing to do with direct messages. This
branch had also written a 5.json, for a different schema. Resolved by
restoring mantra's 5.json untouched and moving the direct message columns
to an AutoMigration(7, 8) with a regenerated 8.json. Taking either 5.json
over the other would have left every device validating a migration chain
against a schema it was never built from; keeping version = 5 would have
made a v7 install refuse to open at all.

The regenerated 8.json is two ADD COLUMNs and nothing else, same as before.

fromGroupEventResult was restructured on mantra: the kind switch moved into
applyInnerEvent, and a SubmissionEvent envelope now wraps nip30303 payloads.
Took that structure and re-applied the direct message branch ahead of it
rather than inside it -- a gift wrap is not a nip30303 payload to apply, and
what happens to it depends only on whether this device's key opens it, so it
does not belong in a function about applying submissions.

The isUserMessage fix was re-applied to the eight call sites mantra's
version has, up from the six it had here.

ChatMessageListViewModel and ChatRoomMessagingScreen took mantra's versions
with the composer state, the two renderings and the reply action layered
back on.

docs/README.md keeps both new rows and mantra's closing note about the
skipped-keys document.

108 tests pass, up from 50 here and 83 on mantra.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 23:45:08 +02:00
Kgothatso Ngako
a909108300 feat: announce which key a room signs with, instead of rederiving it
A signer holds a different secret share under every ceremony it took part
in, and signing with the wrong one produces a partial signature that
cannot aggregate. Nothing said which was which: FrostSigningManager
found a room's key by walking every ceremony this device holds a share
for and rederiving each one's room id until one matched.

That search can only find rooms derived at the one path the constant
names. SharedKeyDerivation.parsePath was written to lift that limit and
was never called, so a room derived anywhere else was invisible to
signing.

So the coordinator now says it. GroupKeyStateEvent (kind 30326) carries
the threshold public key, the ceremony that made it and the path the
room's id came from, posted into the room as its first application
message and filed as a GroupKeyState row. completedKey reads that row
first and follows it to the share.

Nothing secret travels. Every member of the room can read the event, so
a share on it would be each member holding everyone else's -- a 1-of-n
key wearing a t-of-n's clothes. The event names the ceremony; the share
stays in DkgSession.secretShare on the device that generated it.

The coordinator is untrusted, as everywhere else in the ceremony, so a
state is verified rather than believed: the room's id *is* the threshold
key derived at the path, and one that does not rederive its own room is
dropped. That is the same guarantee the rederivation gave, kept rather
than traded for a lookup. The old scan stays behind it for rooms that
predate the table.

Announced after the members are added, which is the only order that
works -- adding them commits a new epoch and MLS will not let a member
read what was encrypted before the one they joined at. A member invited
later still misses it and falls back to the scan, which is where every
member was before this existed.

Replacement is this app's job. These are rumors inside a Marmot group
event, so no relay applies the 3xxxx rule, and the DAO keeps the newest
announcement per room so a backfill cannot walk a room backwards.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 23:36:57 +02:00
Kgothatso Ngako
f5eb744ca7 test: cover the long-running sync, and open the seams needed to do it
The six commits that built the live chat sync added no tests. Everything they
touch fails silently by nature — a filter that drops messages, a subscription
that stops being replayed, a group whose id never reaches the `#h` tag — so the
symptom is always "some messages didn't arrive", days later, on someone else's
phone. 46 tests, in four files.

**What is covered**

RelayPoolSubscriptionTest (13) — the pool's half of surviving a dropped socket.
A query is retained and replayed on reconnect; a closed one is forgotten and
stops the socket reconnecting for it; closing one of two leaves the other alone;
a negentropy exchange is never replayed (its rounds are stateful, so resuming
one reconciles against a conversation the relay is no longer having); an update
to a live subscription replaces what gets replayed, including when the send
itself fails; dropping a relay or closing the pool forgets what they carried;
replay is scoped to the relay that reconnected. Plus the semantic the whole
change rests on, asserted in both directions: a live subscription keeps
delivering after EOSE, a one-shot query still ends at it.

LiveSubscriptionReconcileTest (12) — the requirement this all exists for: the
group filter follows group membership with nobody calling a subscribe function.
Joining widens the filter *in place* rather than reopening (a reopen would drop
the live tail of every other group in that chunk); leaving drops one; leaving
everything closes the subscription; churn inside the debounce window collapses
to one update; a NIP-17 room never becomes a group subscription. Then the
collect loop: events stored against the relay they came from, an event after
EOSE still stored, a CLOSED reopened once the back-off elapses and not before,
and a rate-limited CLOSED waiting far longer — but still coming back.
Backgrounding closes and foregrounding rebuilds, reconnects, and queues the
catch-up.

LiveSubscriptionPlanTest (11) — the filter and planning rules, led by the one
most likely to be "tidied up" later: the gift wrap filter carries no `since`,
because NIP-59 randomizes created_at into the past and a `since` near the
present silently drops new messages.

RelayBackPressureTest (4) and ReconnectBackoffTest (6) — the two pure decisions.
Which CLOSED reasons mean "ease off", and the backoff arithmetic including the
exponent clamp: 2.0.pow(4000) is Infinity and Duration * Double throws on it, so
without it a socket failing long enough turned its reconnect loop into a crash
loop, at the point the network was least likely to recover unaided.

**Seams opened to get there**, each a readability win on its own terms:

  - NostrSocketClientFactory becomes an interface with DefaultNostrSocketClientFactory
    behind it, so the pool can be driven by a fake socket.
  - RelayPool takes its CoroutineScope, so the replay a reconnect triggers can be
    observed rather than raced.
  - LiveSubscriptionManager depends on a new LiveSubscriptionTransport (4
    methods) rather than RelaysSocketManager, which observes the active wallet in
    its init and cannot be stood up in a test at all.
  - Its pure planning helpers move to the companion as `internal`, and its
    launches inherit the caller's dispatcher instead of pinning Dispatchers.IO.
    SynchronizationViewModel already launches observe() on IO, so nothing moves —
    but a coroutine that picks its own dispatcher cannot be driven by a test
    scheduler.
  - reconnectDelay is extracted to ReconnectBackoff.kt with jitter as a
    parameter, so the arithmetic can be pinned without randomness.
  - endsLiveSubscription names the live-subscription termination rule next to
    isTerminalFor, which is the one-shot rule. Having both named makes the
    difference between them reviewable rather than implicit.

kotlinx-coroutines-test is added to commonTest: the pool's bookkeeping is all
suspend functions and there is no runBlocking in a common source set.

The tests were checked by mutation, not just by passing — reintroducing a
`since`, making EOSE terminal, dropping the leftGroupAt filter and removing
retention from query() each produce failures.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 23:31:11 +02:00
Kgothatso Ngako
0319f1613b Merge branch 'mantra' into claude/nostr-event-save-issue-6e9467 2026-09-05 23:30:28 +02:00
Kgothatso Ngako
5321e4af72 Merge branch 'mantra' into claude/distracted-franklin-e95ba4 2026-09-05 23:24:45 +02:00
Kgothatso Ngako
fb21678813 test: pin where a commit's bytes land when the row recording it is written
The mis-routed `framedCommitBytes` fixed in the previous commit was invisible for
one reason: nothing anywhere covered the persisted row. The bytes that reach a
relay come off the in-memory `CommitResult`, so the wire path stayed correct and
the stored path was wrong, and no test looked at the stored path.

## Why the mapping moved before it could be tested

A test that built `MarmotCommitResult` itself would have been writing its own copy
of the mapping and asserting against that. It would have passed against the buggy
code, because the bug was at the call site the test was not using.

So the mapping is now `MarmotCommitResult.from`, called by
`MarmotOutboundDao.inviteMember` and exercised directly by the test. That also
removes the shape that produced the bug rather than just the instance of it: the
old call site listed its named arguments in an order different from the
declaration, which is what put `preCommitExporterSecret` and `framedCommitBytes`
two lines apart. `from` lists the payload in declaration order, in one place, so
there is no second site to get wrong.

## What is covered

Four tests, each payload given a distinct self-identifying value so that a field
arriving in the wrong column names both halves of the mistake instead of comparing
equal by accident:

  - every payload field lands in its own column.
  - the framed commit column never holds the exporter secret -- the regression,
    stated as an invariant rather than an equality so it keeps holding for a
    `CommitResult` this test did not anticipate.
  - a `CommitResult` that never framed its commit still stores a commit. quartz
    defaults `framedCommitBytes` to `commitBytes` and the entity repeats that
    default; the fallback must not quietly become the secret either.
  - the bookkeeping `DatabaseNostrRepository` reads back on acknowledgement is
    carried through. `id`, `chatRoomId`, `userPublicKey` and
    `peerKeyPackageEventId` are all 64-char hex, so two of them swapped in `from`
    would typecheck exactly as silently as the original bug.

Checked by reintroducing `framedCommitBytes = commitResult.preCommitExporterSecret`
into `from`: three of the four fail. A green suite that would stay green against
the bug it names is not coverage.

## What is not covered, and why

That the bytes published equal the bytes stored -- the property one level above
this one -- still is not. It needs the DAO, and the DAO needs Room: `commonTest`
carries only `kotlin.test`, the room3 KSP processor is registered for the android
and ios targets alone with `kspJvm` commented out, and `getInMemoryDatabaseBuilder`
wants a `PlatformContext` no unit test has. That is a Robolectric or instrumented
target, which is a larger change than this fix earns and is better decided on its
own merits than smuggled in here.

The ack-triggered rebroadcast that would have turned the bug into a live fault does
not exist yet, so there is nothing to test there either. When it is written, the
invariant it needs is already asserted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 23:24:21 +02:00
Kgothatso Ngako
ad3304a665 refactor: build the DM inbox filter once, where it can be asserted
The filter fix a commit ago changed a value inline in a ViewModel, which is
not a place a test can reach: ChatMessageListViewModel needs a repository and
a coroutine scope to construct, and NostrDao needs Room. So the filter that
had just been wrong in three call sites went back to having no coverage at
all.

Nip17Filters.inbox is that filter with one definition. ChatMessageListViewModel
and ChatRoomListViewModel now both call it — they had been building it
separately and identically, which is also what made their negentropy requests
collapse into one under computeId, a coincidence better expressed as shared
code than left to hold by luck.

Nip17FiltersTest asserts every clause that was got wrong in production:

  - the p tag names us, not a peer
  - there is no authors clause, because a wrap is signed by the throwaway key
    GiftWrapEvent.create mints and discards, so authors=[anything knowable]
    matches nothing on any relay
  - there is no since cursor, because NIP-59 back-dates a wrap by up to two
    days and a high-water mark taken from the newest wrap we hold skips mail
    stamped behind it — the trap waiting for whoever acts on the TODO in
    NegentropySynchronizeRequest.toSynchronizeNostrEventRequest
  - the wire JSON is pinned, so an added default cannot quietly split the two
    callers back into separate requests
  - the SQL NostrEventFilterQuery builds from it bounds no author either,
    since negentropy is only as good as the agreement between the set we build
    locally and the set the relay builds from the same filter

Neither of the two failure modes this covers was visible from reading the
filter. The authors clause failed silently for as long as it existed, and the
peer p-tag failed loudly but somewhere else entirely — in a Room transaction,
three files away, as a MAC error out of Nip44.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 23:22:06 +02:00
Kgothatso Ngako
e1d35bbd6c test: pin who can open a gift wrap, and what happens to everyone else's
The Invalid Mac crash had no test standing between it and a repeat, so this
adds one that reproduces it.

GiftWrapMessageTest builds real NIP-59 wraps with real secp256k1 rather than
recorded fixtures. The property under test is the key agreement itself —
whether ECDH(ourPriv, ephemeralPub) can stand in for the conversation key the
wrap was sealed under — and a fixture would only prove that the fixture still
parses. Three cases carry the regression:

  - someone else's mail comes back null rather than throwing
  - not even the sender can reopen what they sent
  - isAddressedTo answers exactly what unsealing would

Checked against the reverted fix, those three fail with the production
exception verbatim (java.lang.IllegalStateException: Invalid Mac: Calculated
bf2e6480…), while the two describing behaviour that never broke — the happy
path, and isAddressedTo's reading of the p tag — stay green. A test that
cannot fail against the bug it names is not worth the run time, so the split
matters.

The last of the three is the one guarding the fix's structure rather than its
outcome. NostrDao decides whether to index on isAddressedTo, then throws
GiftWrapUnsealException if decryptGiftWrapSeal returns null anyway; those two
answers have to agree for either path to be correct. If they drift, the DAO
either skips mail we can open or resumes rolling back transactions, and
neither shows up as a failure anywhere near the change that caused it.

commonTest gains kotlinx-coroutines-test for runTest. decryptGiftWrapSeal is
suspending, runBlocking does not exist in common code, and every layer worth
testing below the ViewModels — DAOs, repositories, the model's crypto — is
suspending too, so the dependency pays for more than this file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 23:21:52 +02:00
Kgothatso Ngako
42dd38cfc4 test: pin the two invariants this session left unguarded
Both are silent when broken, which is why they are worth asserting rather
than reasoning about.

**The cache's reuse decision.** MlsGroupCache exists because quartz drops
a secret tree's skipped-generation keys on save, so rebuilding a group
between two messages loses any that arrives late. Its safety argument is
one comparison: reuse while the stored state is still what the cache last
wrote, rebuild when it is not. Get that wrong in either direction and
nothing complains -- reuse too eagerly and a group carries on from a
ratchet another writer already moved, which corrupts decryption rather
than failing it; reuse too rarely and the cache does nothing and the
original bug is back with no symptom.

That decision is now a generic LiveInstanceCache with MlsGroupCache as a
typed facade over it, so it can be tested without standing up an MLS
group. Splitting it also made two behaviours explicit that were previously
incidental: a failed build no longer leaves the old instance behind, and
an instance whose use threw is deliberately not cached -- it is
half-advanced and never persisted, so the next caller has to start from
disk.

**Rumor and row ids agreeing.** MantraDao writes an entity whose id comes
from fromXEventTemplate and separately builds the rumor it submits with
rumorOf, which hashes the template itself. Both are meant to produce one
id and nothing checked it. Diverging would mean submissions naming an
event nobody has, deleteByPayloadEventId silently un-queuing nothing so
superseded translations go out anyway, and every receiver creating a
second row instead of converging on the sender's -- all of it invisible,
since the ids are opaque hex either way. Asserted per kind, plus the whole
chain out through the submission envelope.

Both suites were mutation-checked rather than trusted: inverting the
staleness comparison fails one cache test, recording the pre-block state
fails another, and hashing the rumor under a different author fails all
six id tests.

Still uncovered, and not cheaply fixable: FrostSigningManager's and
MantraDao's state machines both need a Room harness, and commonTest has
none. The FROST crypto path is covered by FrostSigningRoundTest; the
message-driven parts around it are not.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 23:18:54 +02:00
Kgothatso Ngako
a74a4b71cf test: cover the two decisions that decide who said what
The crypto was tested; the logic that acts on it was not. Both untested
pieces were the security-critical ones, and neither fails loudly when it
goes wrong -- one silently widens who may impersonate whom, the other
silently destroys a message.

Extracted MarmotDirectMessage.classify, which decides what an arriving
wrap is to this device, from ChatMessage.directMessage, which turns that
decision into rows. The decision is pure; only the filing needs a
database, and Room-backed code cannot be unit-tested in this project. Same
split, and for the same reason, as pulling the wrap/open crypto out of the
DAO in the first place.

Extracted MarmotInboundManager.mip03Rejection for the same reason. Its
kind:1059 exemption is the most dangerous line in this feature: widened to
another kind, or stripped of its kind guard, it hands every member of
every group the ability to publish events as anybody, and nothing else in
the pipeline would notice. There is now a test that walks seven kinds and
asserts each is still held to MIP-03.

Fifteen cases, the ones worth naming:

`our own message is ours, even though we cannot open it` and `ours is
decided before anything is opened`. A sender cannot decrypt their own wrap
-- the key was discarded -- so by decryption alone this is
indistinguishable from a bystander's view, and only the MLS identity
separates them. Get it wrong and the inbound path files an empty
placeholder over the row sendChatMessage wrote, which holds the only copy
of those words. It is the one failure here that loses data rather than
rendering something wrong.

`words sealed by one member and sent by another are dropped`. The check
that replaces MIP-03 for this kind, tested directly rather than described
in a comment as it was before.

One test asserts something I had wrong. I expected a seal relabelled with
another member's pubkey to be caught by the signature check; it never
reaches it. NIP-44 derives the conversation key from the pubkey being
claimed, so relabelling a seal makes it undecryptable by the person it was
encrypted for -- the label is bound to the key, not merely asserted
alongside it. The outcome is Unreadable, which is the truth: the recipient
genuinely cannot read it. `a seal tampered with after signing is dropped`
covers what verify() does catch, using an alteration that survives
decryption.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 23:17:43 +02:00
Kgothatso Ngako
b4ac65f5c9 feat: sign a nostr event with the group's shared key
A ceremony leaves every member holding a share of a t-of-n key and no way
to use it. This is the other half: a session that turns an unsigned nostr
event into one signed by the group.

The shape is ChillDkgRitualManager's, deliberately. The member who
proposes coordinates, protocol messages travel as gift-wrapped rumors on
the same NIP-17 pipeline chat messages use, each inbound message is
persisted and then the session is asked whether it can move, and every
step is recomputed from stored inputs so a device killed mid-round
resumes on the next message. Anyone who has read that manager can read
this one.

    proposer --[ 30320 proposal   ]-> everyone   the unsigned event
    signer   --[ 30321 nonce      ]-> everyone   this device's public nonce
    proposer --[ 30322 signer set ]-> everyone   who signs, and their aggregated nonce
    signer   --[ 30323 partial    ]-> everyone   this device's partial signature
    proposer --[ 30324 signature  ]-> everyone   the finished 64-byte signature
    anyone   --[ 30325 failure    ]-> everyone   abandon + blame

Three things are genuinely different, and each is why this is a separate
manager rather than another branch of that one.

**It does not need everybody.** A DKG cannot finish until every member
takes part; that is what makes the key. Signing needs t, and waiting for
n would throw away the property the group ran a ceremony to get. So the
coordinator waits for the threshold to be reachable, picks a set and says
who is in it. Members left out do nothing and stall nothing.

**Restart-safety is forced rather than chosen.** SecretNonce cannot be
serialised and refuses to be used twice, so storing the randomness it
derives from and regenerating on demand is the only way a session
survives the app closing. That is safe for exactly one reason: a session
signs one message and cannot be made to sign another. Two rules hold it
in place and both are load-bearing rather than tidy:

  - the event id is written at creation, and a proposal that disagrees
    with it is refused rather than applied;
  - the aggregated nonce and signer set are write-once. A coordinator
    that sends a second, different set is ignored. Obeying it would mean
    two partial signatures over one secret nonce against two challenges,
    which is precisely how a secret share is extracted. The session
    stalls; the share does not.

**One approval, not three.** A DKG asks three times because each step
publishes something different and commits the member to something
different. Here every step serves one decision -- sign this event or do
not -- and the event is fixed before the member is asked, so a second
prompt would be the same question twice. Declining is broadcast rather
than silent: a t-of-n group can sign without you, but only if it knows.

Two things are checked rather than trusted, both because the coordinator
is untrusted by construction: the event id is recomputed from the
proposal's own fields, so a proposer cannot have the group sign one thing
while showing them another; and the finished signature is verified before
the session is called complete, so a bad aggregate is a failure here
rather than a rejection at every relay it reaches.

Signer ids are derived, not stored: a member's FROST id is their index in
the bytewise sort of the ceremony's host keys, the same ordering ChillDKG
hashed into the session identity and the same one the public shares are
in. Deriving means signing cannot disagree with the ceremony that made
the key.

DkgSession gains publicShares, kept because FROST validates each signer's
secret share against its public one. A ceremony finished before this
column reads back null and signing runs without that check rather than
refusing.

The tests run the same calls in the same order against real FROST and
assert the aggregate verifies as a nostr signature. That path was written
from reading the library rather than from a working example, so it is the
part most likely to be subtly wrong -- and wired up wrong it fails
silently, on every device.

Kinds start at 30320 with a gap. The DKG runs 30310-30316 and the
nip30303 document kinds run 30300 up; those two already collide at 30310
and 30311, and SubmissionEvent sits on 30312, which is also the DKG's
round-1 kind. They are kept apart today only by riding different
transports, which is luck. Signing shares a transport and rooms with the
DKG, so it starts clear of both.

No UI yet: this is the session logic, reachable through proposeSigning,
approve and decline.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 21:38:23 +02:00
Kgothatso Ngako
fcc28de931 Revert "fix: hold a payload whose parent has not arrived instead of losing the event"
This reverts commit d7aac49.

Reverting restores the defect it addressed: a payload referencing a row
the receiver does not have violates a foreign key, and SQLite aborts,
rolling back the whole inbound transaction -- the nostr event, the group
event, the submission and the transcript line, none of them retried.
That is what produced the observed `FOREIGN KEY constraint failed` on an
artifact whose dialect had not arrived.

Also drops the schema back to v5. Any device already migrated to v6 will
refuse to open its database, since the builder sets no destructive
fallback on downgrade; clear that app's data before installing a build
from this commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 21:12:50 +02:00
Kgothatso Ngako
d7aac49cf1 fix: hold a payload whose parent has not arrived instead of losing the event
A receiver hit `FOREIGN KEY constraint failed` on an artifact submission
and lost the whole group event. The artifact referenced a dialect the
receiver did not have, MantraArtifact.dialectId is a foreign key, and
SQLite answers a violated constraint by aborting -- which rolled back
the entire transaction the inbound pipeline runs in. Gone with it: the
NostrEvent, the MarmotGroupEvent, the submission's MarmotInnerEvent
holding the payload verbatim, and the transcript line. Nothing retries,
so the artifact stayed lost even once the dialect turned up.

Every nip30303 entity is a child of another and the schema enforces all
of it -- artifact→dialect, version→artifact, chapter→version,
chunk→chapter, translations→both of theirs -- so this was every branch,
not one.

And submissions make arriving before your parent ordinary rather than
exotic. That is the point of them: an admin submits a backlog in
whatever order they hold it, and a member who joined last week can be
sent what the group was told last month. Both produce payloads whose
parents are not here yet, and both were losing data.

So check the parents before inserting. A payload that arrives early is
held on the submission row -- awaitingEventId names what it waits for --
and applied when that arrives. Releasing one can release another, a
version freeing its chapters and those freeing their chunks, so it walks
outward until nothing more comes unstuck. A payload with a second parent
still missing is re-pointed at that one rather than retried on every
arrival.

Nothing is written to the transcript while a payload is held. Nobody has
said anything yet; the line appears when it is applied, in the position
its own timestamp gives it.

Two things fall out of the shape:

parentRefsOf is pure and separate from the lookups, because the mapping
is the part that can silently drift from the schema and there is no
database harness in commonTest to catch it. ParentRefsTest pins one case
per kind. Which table an id lives in is carried as the kind of event
that would have created it, so there is no second enum to keep in step.

applyInnerEvent takes ids rather than a GroupEvent, since replay happens
long after that object is gone. A released payload is recorded as not
ours: we hold the parents of anything we wrote, having written those too.

Also reconstructs a held bare nip30303 event from its own columns rather
than parsing its content as an event -- only submissions carry an event
there, and reading both that way would have stranded every bare one
permanently.

Verified: the v5→v6 migration runs clean on the receiver's real
populated database. The hold path itself still needs a fresh submission
from a sender to exercise end to end.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 21:05:35 +02:00
Kgothatso Ngako
fa380e94e1 feat: add a SubmissionEvent that carries a nip30303 event as its payload
Every nip30303 kind so far describes a thing: an artifact, a dialect, a
chapter, a translated chunk. None of them describes the act of putting
one in front of a group, and until now nothing needed to -- a group
event's sender was the author of the event inside it, so the two
questions had one answer by construction.

That construction is also the limit. It means a group can only ever hold
work written by its own members under their own keys. A translation
lifted from a public archive, a chapter transcribed by an outside
contributor, an artifact somebody published years ago: none of it can go
in without a member re-authoring it and taking the byline.

Kind 30312 is the envelope that separates them. Its content is the
payload event's JSON, whole -- same id, same pubKey, same signature,
nothing rewritten to look like the submitter's work. The submitter signs
for the envelope; the author still signs for the event. Two tags name
what is inside so a client can decide whether it can apply a submission
without parsing the content first:

  payloadKind   the payload's kind
  payloadId     the payload's id, with the author slot carrying the
                payload's author -- who, unusually for an id tag in
                this package, is often not the event's sender

Kinds 30300-30311 are taken (30305 and 30307 by contributor lists), so
30312 is the next free one.

A submission is not an endorsement and grants nothing. Who may submit is
the group's business; this only makes the question expressible.

The test covers the property the whole thing rests on: an event written
by an outsider goes into an envelope, comes out of a JSON round trip
with its id, author and signature intact, and does not acquire the
submitter as its author. It also pins payload() returning null rather
than something empty when the content will not parse -- which needed
android.util.Log stubbing, since quartz logs on that path and unmocked
Log methods throw, failing the test on the log line rather than on what
it came to check.

Nothing sends or reads one yet.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 20:21:00 +02:00
Kgothatso Ngako
79e99ae702 feat: build the envelope a direct message travels in
A one-to-one message inside a Marmot group is a stock NIP-59 gift wrap
carried as the MLS application payload: a throwaway-keyed kind:1059 around
a sender-signed kind:13 seal around the kind:14 rumor holding the words.
Every member decrypts the MLS layer and sees the wrap; only the recipient
can open it. See docs/marmot-direct-messages.md.

This is the crypto on its own, with no database and no MLS state, because
the outbound path (the notary) and the inbound path (the kind switch in
ChatMessage) both need it and neither can be unit-tested -- there is no
sqlite driver on the JVM test classpath. Extracting it first is what makes
the ten tests here possible; real secp256k1 does load under
testDebugUnitTest, so none of this is mocked.

Three choices worth stating, all of them consequences of the wrap using a
throwaway key rather than the sender's own:

Nothing in the wrap names the sender. GiftWrapEvent.create mints and
discards its own random key, so who sent a message comes from the MLS
frame around it -- authenticated to a leaf, and unforgeable -- rather than
from a self-asserted pubkey field. The seal inside is the only layer the
sender signs, which is what the inbound path will bind to the MLS sender
identity before it renders a word.

The sender cannot reopen their own message. The throwaway key is gone at
send time and nothing reconstructs it. `the sender cannot reopen their own
message` asserts that rather than leaving it to be discovered, because the
obvious fix -- persisting the throwaway private key -- would be strictly
worse than the identity-keyed wrap this was chosen over, and would
reintroduce the attribution the throwaway key exists to remove.

No layer is fuzzed. NIP-59 randomises the wrap and the seal by up to two
days to frustrate correlation at a relay, and both GiftWrapEvent.create
and SealedRumorEvent.create default to it. There is no relay at this layer
and the kind:445 already carries the true time, so fuzzing would only
scatter the "sent a private message" line up to two days out of position
in every other member's transcript.

open() returns null rather than throwing on every way a wrap can fail to
open -- somebody else's message, a malformed payload, a layer that is not
the kind it claims. Its caller is midway through processing a kind:445
that may carry a perfectly good message for somebody else, and an
exception would abandon all of it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 18:57:11 +02:00
Kgothatso Ngako
9f14679aac feat: let the coordinator open a #admins room keyed on the shared key
Once a ceremony completes, the shared-key screen offers its coordinator a Marmot
room named "<group> (#admins)" with every member of the ceremony in
MarmotGroupData.adminPubkeys. The room the ceremony ran in is NIP-17, where nobody
administers anything; this gives the same people a room where every one of them
can act, which is the shape a group that has just made a t-of-n key is asking for.

Built directly rather than through MarmotGroupData.bootstrap, which hardcodes a
single admin, and baked into the epoch-0 GroupContext so later invitees receive a
populated group from their welcome instead of chasing a bootstrap commit that
predates their membership.

## The id is derived, not random

Every other Marmot room mints `nostrGroupId` as RandomInstance.bytes(32). This one
derives it from the group's threshold key, settling the
`// TODO: Generate GID through frost...` already sitting in
SelectChatRoomTypeViewModel.

Derivation buys two things random cannot. Every member's device can compute the id
from a ceremony they all took part in, so the room is addressable without being
announced; and two members racing to create it arrive at the same id rather than
two rival rooms -- which is why createAdminGroup returns to the existing room
instead of minting a second one.

## Why the derivation is what it is

SharedKeyDerivation walks the path as successive FROST tweaks, one per index,
returning both the XonlyPublicKey and the TweakCache. The cache is not an
optimisation: a signing session created without the same tweaks aggregates to
signatures that verify against a different key, which is why the id is usable as
an identity later rather than only as a label.

It is not BIP32, and the doc comment argues that at length rather than leaving it
to be rediscovered. A BIP32 node is a key *and* a chain code; ChillDKG produces no
chain code. BIP32 wants one only because it computes the tweak scalar for you, and
a FROST tweak takes that scalar as an input -- so choosing it directly removes the
chain code from the problem rather than requiring one to be invented and agreed
forever. It also removes a trap: with x-only keys there is no single obvious
serP(K_par), and two devices picking different parity conventions would silently
derive different keys rather than fail.

Each scalar commits to the key being tweaked as well as the index, so steps cannot
be reordered or replayed at a different depth. Tests cover that, determinism
across calls, path and key sensitivity, and that the cache and the public key
agree.

Hardened derivation is not available here and never will be: it needs the parent
private key, which in a threshold group nobody has. That leaves the non-hardened
weakness -- k' = k + t with publicly computable t inverts -- so anyone learning one
derived private key recovers the group key and can sign with no quorum at all. The
rule that follows is stated at the top of the file: never reconstruct a derived key
in the clear.

## The path is recorded in the room

MIP-01's group data is a fixed TLS schema with no extension map, so a custom field
would emit bytes other Marmot clients cannot decode. The path rides in the
description instead, on its own line under a marker, so somebody rewriting the
rest of the description does not cost the group the record of how its key was
derived:

    Admins of Ubuntu Collective.

    Shared key path: m/9420/0/0

Worth storing although the path is currently a constant: it is what rebuilds the
TweakCache a signing session needs, and recomputing from the constant only holds
while the constant never changes. parsePath refuses hardened indices rather than
tolerating them -- such a path cannot have been walked here, so acting on one
would derive something other than what the room claims.

## Known limits

Members without a published MarmotKeyPackage cannot be invited; inviteAdmins
collects them and logs them, and the coordinator is not yet told.

Invites go one at a time, each advancing the MLS epoch, so the room is re-read
between them. That inherits a silent failure mode documented in
docs/marmot-membership.md: the first invite takes the deferred-welcome path even
though the group is still just its creator, and a commit reaching a member before
their welcome is dropped rather than queued. Not introduced here -- group creation
has always done this -- but more visible in a room whose whole membership is known
up front.

Nothing here has run on a device.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 14:38:58 +02:00
Kgothatso Ngako
661a5caa17 fix: build the local negentropy set from the whole filter, not a guess at its shape
A negentropy exchange compares two sets defined by the SAME filter: the relay
builds its side from the filter carried in NEG-OPEN, and this device builds its
side from getNegentropicNostrFeedIds. Any clause we fail to apply locally makes
our set a superset of the relay's, and each extra row comes back as an id the
relay is "missing" -- which this app then queues as a broadcast. Any clause we
apply more tightly makes it a subset, and the difference comes back as ids to
re-download that we already hold. Neither shows up as an error; both show up as a
sync that never settles.

getNegentropicNostrFeedIds was a `when` over the shape of the filter, dispatching
to one of eight hand-written @Query methods. Each method could only bind the
parameters it happened to declare, so the branches disagreed with the filter they
were serving:

  - `until` was expressible by NO branch. It is sent to the relay in NEG-OPEN and
    was never applied here, so every local event past the requested window was
    reported to the relay as one it lacked.
  - `since` was strict (`createdAt > :since`) where NIP-01 is inclusive, so an
    event stamped exactly on the boundary was a phantom "need" on every pass.
  - `kinds && authors` was tested before any tag branch, so a filter carrying
    kinds, authors AND tags silently dropped the tags. `kinds && ids` dropped
    authors. Every branch dropped whatever it had no parameter for.
  - tags were matched with `tags LIKE '%' || :value || '%'` -- a substring scan of
    the serialized tag JSON that matches the value in ANY tag position. A pubkey
    referenced in an `e` tag counted as a `p` match. And only `tags[name].first()`
    was ever bound, so the second and later values of a tag were dropped.
  - the reply branch matched `'%' || :eventId || '%reply%'`, which needs the
    literal text "reply" to appear somewhere after the id: it misses
    `["e","<id>"]` with no marker and false-positives on any later tag containing
    the word.
  - the `else` branch ignored the filter's kinds entirely and substituted
    `arrayOf(TextNoteEvent.KIND)`. A filter with only authors, or only tags, got a
    local set of kind-1 notes -- unrelated to what the relay was reconciling.
  - more than one filter returned emptyList() with a "not yet supported" warning.
    That is the worst available answer: an empty local set tells the relay we hold
    none of these events, so it hands back its entire set as ids to download.
  - the limit branches ordered `createdAt ASC LIMIT n`, returning the OLDEST n
    where a relay answering a limited filter returns the newest.

## The replacement

NostrEventFilterQuery translates a SynchronizationFilter into one SQL statement
that applies every clause, and NostrEventDao.getNostrEventsMatchingFilter runs it
as a @RawQuery. Raw because a nostr filter is a variable set of constraints over
variable-length lists, which is precisely what @Query cannot express -- and what
drove the per-shape methods that dropped constraints in the first place.

Semantics follow quartz's FilterMatcher, which is what the relays this app talks
to implement: membership for ids/authors/kinds; AND between tag names and OR
between the values of one name for `tags`; AND both ways for `tagsAll`; inclusive
`since`/`until`; and a present-but-empty list matches nothing.

Tags are matched by looking for the `["<name>","<value>"` fragment, built by
encoding through the same serializer that wrote the column so escaping agrees,
with `%`/`_`/`\` escaped and `ESCAPE '\'` on the LIKE so a wildcard inside a value
cannot widen the match. Anchoring on the tag name and on the closing quote of the
value is what keeps a hex string from matching in an unrelated tag position.

Multiple filters are now the union of their matches, de-duplicated by id.

## The Marmot branch is kept, and narrowed

Group messages still answer from MarmotGroupEvent: that table carries the NIP-40
expiry a relay uses to decide whether it still serves an event, and an indexed
chatRoomId instead of a scan of the tags JSON. But the branch now only claims a
filter it can fully honour -- exactly kind 445, an `h` tag, and nothing else --
because it answers from a different table and would otherwise reproduce the same
silently-dropped-constraint bug it is an exception to. It also fills in the `h`
tag and the real signature on the NostrEvent it synthesizes rather than leaving
them empty.

## Tests

NostrEventFilterQueryTest pins the generated SQL and the bound values for each
clause, including tag escaping and the empty-list case. It asserts the
translation rather than eyeballing it, because a dropped clause is not an error
at runtime -- it is reconciliation quietly reporting differences that are not
real.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 11:25:15 +02:00
Kgothatso Ngako
61869f0046 test: run a real ChillDKG ceremony through the ritual's ordering rules
ChillDKG has no session-params object the group agrees on out of band: every step
takes the host public keys and the threshold and hashes them into the session
identity itself. A group whose devices order their participants differently
therefore gets no key at all, and nothing in the protocol tells you that is what
went wrong. ChillDkgRitualManager has each device derive that order
independently -- sort the collected host keys, and order each round's messages by
their sender's host key to match -- and until now nothing checked that the two
rules agree, or that they agree with what ChillDKG expects.

Four tests, against the real library rather than a stand-in:

  a ritual ordered by host key produces one shared key
      A full 2-of-3 run -- step1, coordinatorStep1, step2, coordinatorFinalize,
      participantFinalize -- with the participant set built by hostPublicKeys()'s
      rule and both rounds ordered by orderedPayloads()' rule. Asserts every
      member lands on the same threshold public key and on distinct shares.

  sorted host keys give every device the same participant order
      The same members in three arrival orders, since relays deliver host keys in
      whatever order they please, must sort to one order.

  one device ordering participants differently gets no key
      The negative that keeps the other two honest: with one member running the
      same people in another order, some step has to fault. Without this a broken
      ordering rule could pass the happy-path test by being uniformly broken.

  host keys are not the nostr keys they come from
      deriveHostSecretKey's two obligations: it must not hand ChillDKG the nostr
      identity key (a flaw in either protocol would otherwise reach the other),
      and it must be deterministic, or a reinstall cannot recover the share.

These live in commonTest and run under `./gradlew :composeApp:testDebugUnitTest`.
The secp256k1 natives do load there: the Android loader fails and falls back to
extracting the JVM platform build, so these are real curve operations, not
mocked ones. Room-backed code still cannot be tested this way, which is why the
manager's database behaviour is not covered here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 00:14:00 +02:00
Kgothatso Ngako
9abdf42921 Refactor torch to mantra 2026-07-15 01:14:46 +02:00
Kgothatso Ngako
bad1b0eb31 Fork Aux to make Torch 2026-06-17 18:10:06 +03:00
Kgothatso Ngako
4a2ddbc4f2 Correct the namespace and introduce an android specific namespace 2026-04-21 23:46:45 +02:00
Kgothatso Ngako
0652c6add4 Pass the torch... initial commit. 2026-03-23 01:41:39 +02:00