Arkus Innovation Studios
Ship Log 001 · The first two weeks at Arkus
A one-off catch-up post. Four mini ship logs covering what Arkus shipped between May 6 and May 20: category-aware evidence scoring in Cipher, recovered extraction in Index, real Live audio in Arkus AI, and a real production bug we found and fixed in VentureIP.
This is a one-off catch-up post. Normally Ship Logs are weekly and single-topic, because that is the format that respects your time and forces me to actually say something. Today, four mini logs in one place, because we are starting two weeks into a cycle and there is work worth covering.
Onward.
Cipher: category-aware evidence scoring
For most of Cipher's lifetime, the implicit standard for "proof" was SaaS-style commercial traction. ARR, paying customers, contracted pipeline. That is the evidence model the broader VC ecosystem reaches for, and we inherited it.
The problem: it under-scored everything that does not prove itself the SaaS way.
A pre-revenue biotech with a Phase II trial reading out next quarter is not a weaker bet than a SaaS company with $40K MRR. A hardware company with a signed manufacturing partner is not weaker than a content platform with a paid trial. Cipher was treating them as weaker, because revenue is what its scoring math knew how to count.
We shipped a six-family evidence taxonomy: commercial, technical and IP, regulatory, partnership, operational, and qualified government. Pre-revenue companies no longer get penalized across non-revenue dimensions when category-appropriate evidence exists. Anti-inflation guardrails went in for the usual offenders: advisor logos without formal roles, accelerator participation framed as validation, unsigned conversations described as partnerships.
The model was treating "no revenue yet" as "no proof yet." Those are different things.
Index: extraction only counts if it reaches the engine
Index had been drifting toward a fragile pattern. PDFs were getting uploaded, but extraction was sometimes routing into a lightweight mini-report path that bypassed the deep analysis engine. The user saw a "report." It was not the real one.
We restored a text-first extraction pipeline. PDFs now get converted to page-anchored selectable text with PDF.js before Gemini 2.5 Pro sees them. Text is sent as text parts, not base64 blobs. Gemini 2.5 Pro stays the primary deep extraction model. Gemini 3.1 Pro Preview is the analysis-pass fallback. The Gemini proxy now requires authenticated Bearer JWTs. The public Vite Google API key is out of the production bundles entirely.
Currency, budget, and pathway parsing also got hardened. Previously the system was reading words like "base" as billion-scale suffixes. That is funny until you realize a real partner saw a report with that error in it.
Extraction success does not count if it does not reach the analysis engine.
Arkus AI: real Live audio, server-side everything
Arkus AI moved from a readiness UI to a real Live audio runtime. Server-controlled session tokens, reconnect classification, GoAway handling, AudioWorklet streaming on the microphone path, lifecycle states that actually represent what the room is doing.
The constraint we held tight: provider keys stayed server-side. No browser-facing key was added. No transcripts, raw audio, full prompts, or provider outputs get persisted. The Live session endpoint validates session mode and report context before issuing a short-lived credential.
What still needs validation: real provider mic-to-end testing with a configured environment. We have exercised the route with controlled provider failure. We have not exercised it end-to-end with a live mic and a real provider key in the loop.
The boring constraint, held consistently, is what makes a voice feature defensible. The exciting feature is the easy part.
VentureIP: production found a real bug, we fixed it
VentureIP went through the most concentrated hardening this cycle. Returning-member login, consolidated activation through a single redeem path, admin lifecycle controls, email delivery tracking via Resend webhooks (delivered, bounced, complained, opened, clicked), database-backed session reconciliation, same-origin enforcement on cookie-authenticated mutating routes, rate limits on sensitive auth and public endpoints, append-only admin audit events, CI.
The honest part: during live seeded testing, password setup threw a 500 in production. The cause was an activation event write hitting the database before the new user row had committed. Postgres returned a foreign-key error. We fixed it by putting the event write in the same transaction as the user creation, where it should have been to begin with.
That bug only surfaced because we ran the live seeded workflow end-to-end with a real inbox. Local tests were green the whole time.
If you are building activation flows, test the path with a real seed and a real inbox. Mocks lie about the things that matter most.
What is next
Continued reliability work on the document pipeline, particularly claim-1 retrieval on patent-rich submissions where the path needs to be deterministic. Refinement of how regulatory and partner evidence get scored across the lifecycle: engaged, submitted, cleared or signed. End-to-end testing on Live audio and activation with real credentials in the loop.
Back next Friday with a single-topic post. The catch-up format is a one-off.
Sheldon