aiagency.nz
Back to blog
Field Notes22 April 20268 min read

Shipping three voice apps in six months, what we learned

Notes, Finance, Clients. Same engine, three audiences. The bets that paid off and the assumptions we had to throw out.

Shipping three voice apps in six months, what we learned

Building SpeaktoNotes was the starting point. A voice-to-text app with smart formatting, six output templates, and a writing style engine that adapts to how the user actually speaks. It solved a problem we’d seen repeatedly across NZ: tradies, sales reps, real estate agents, all drowning in information they’d captured by hand when they could have spoken it in seconds and moved on. The app launched across iOS, Android, Mac and Windows, and it did what it said.

SpeaktoFinance came next, and it forced a sharper decision about scope. Voice capture for financial records: expenses, receipts, tax memos, mileage logs, invoice notes, with enough structure for accountants to work with directly. The critical boundary we built around: the app processes voice input and structures it into records, but it never touches actual money or banking. That kept the compliance surface manageable and made the problem tractable for a two-person studio.

SpeaktoClients was the third piece, and the one that tied the pipeline together. Voice to client communication: follow-ups, proposals, meeting recaps, cold outreach, onboarding messages. You speak rough notes about a client interaction and the app structures them into something ready to send. Gmail and Outlook integration at launch, calendar event creation from voice, the same cross-device sync architecture running underneath all three.

Three separate apps. Same core engine. Six months from first commit to all three live.

Why build three at once

The obvious question is why build three simultaneously rather than shipping one, learning, and building the next. The honest answer is that they were always designed as a pipeline, and a pipeline is only useful whole.

SpeaktoNotes captures thoughts and information at speed. SpeaktoFinance captures the financial layer of those same interactions. SpeaktoClients closes the loop with the people involved. Running only one or two in isolation left the workflow incomplete, and an incomplete workflow is just a feature, not a product.

The shared architecture made parallel development viable. The voice capture layer, the AI processing, the template engine, the cross-platform sync. These were built once and adapted for each app’s specific output. Six templates per app, fifteen templates across the three products, all variations on the same underlying structure.

What the pipeline actually solves

The professionals who get the most from all three apps are the ones whose work is inherently verbal but whose record-keeping is inherently written. Tradies are the clearest example. They quote jobs on site, log expenses as they go, and need to follow up with clients before the day ends. None of that work happens at a desk, and none of it is naturally suited to typing on a phone mid-job.

SpeaktoNotes means the site visit gets captured in full, spoken in the car on the way to the next job. SpeaktoFinance means the materials receipt gets logged immediately rather than sitting in a glove box until Friday. SpeaktoClients means the follow-up email gets drafted from a voice note before the evening is out.

The AI automation layer is what separates this from basic transcription. Raw voice input is messy: incomplete sentences, changed directions, filler words. The processing cleans and structures that input into something genuinely usable. SpeaktoFinance targets 99%+ accuracy on clean voice input for financial data extraction: dates, amounts, vendors and categories coming through structured and accountant-ready, exportable as PDF or CSV.

The question-first model

One thing that worked consistently across all three apps was structuring the AI around questions rather than commands. Instead of requiring users to speak in a specific format, the apps ask short clarifying questions to fill in what they need. A follow-up email needs a recipient, a context, a next step. A tax memo needs a date, an amount, a category. Rather than front-loading all of that into one voice input, the apps guide the capture through a short question flow.

This made a meaningful difference to output quality. Users speak naturally, the app fills the gaps intelligently, and the structured result is complete without requiring a second pass. It also reduced friction for first-time users who hadn’t learned to pre-organise their thoughts before recording.

The question-first approach informed the design of all three apps: one action per screen, complexity handled in the AI layer rather than pushed onto the interface.

Six months, three live products

The timeline was tight but not chaotic. The shared architecture was the reason it was achievable. Once the core voice processing pipeline was reliable, each subsequent app was primarily a template and integration problem rather than a foundational engineering one.

SpeaktoNotes established the pattern. SpeaktoFinance applied it to a more constrained domain, since financial data extraction tolerates far less ambiguity than general note-taking. SpeaktoClients extended it to professional communication with email and calendar integrations the first two didn’t require.

The integration work for SpeaktoClients added the most complexity in the final stretch. Getting reliable calendar extraction from natural speech, where someone might say “let’s catch up next Thursday afternoon” without specifying a time, required more iteration than any other feature across the three products.

Planned integrations for SpeaktoFinance with Xero and QuickBooks will extend the accountant-ready export into direct sync. That sits in the roadmap rather than the launch feature set: accountants needed to trust the data quality before they’d want it flowing directly into their systems.

Where it sits now

All three apps are live across iOS, Android, Mac and Windows. SpeaktoClients is in early access. The pipeline works as a unit.

For NZ businesses thinking about AI automation as a practical operational tool rather than a future ambition, the SpeaktoSuite is the most direct demonstration of what voice-first AI looks like when it’s designed around a specific workflow. Not a general-purpose tool, not a chatbot. A pipeline built around the way people who are never at a desk actually work.

The question that shaped all three products was the same one: what would this look like if it was built for someone whose most important moments happen away from a screen? Everything else followed from that.

Read next

Call 020 4164 7596Free report