- Role
- Product designer
- Team
- PM + 4 engineers
- Timeline
- 6 months
- Outcome
- $400K ARR, year one
Premise
I was the only designer at Canny, working with a PM and 4 engineers to shape Autopilot: AI-powered capture for customer feedback, pulled from the places teams already work.
Feedback everywhere
Think of Canny like Reddit for product feedback. Users post requests, others upvote and comment, and teams use that signal to decide what to build next.
That works while the volume is manageable. Past a certain scale the backlog becomes a job in itself. Posts to review, duplicates to merge, customers waiting on answers.
Feature requests from Canny's users
And feedback had stopped living in Canny. High-signal requests were showing up in support conversations, sales threads, CRMs, Slack, and app reviews. Canny was supposed to be the source of truth, but half the signal never made it there.
Duplicates multiplied. Important asks got buried in threads nobody had time to comb through. The less teams could trust what was in Canny, the less useful it became.
















Feedback, everywhere, all at once
The bet
Instead of asking users to bring feedback to Canny, we wanted Canny to be the place feedback naturally ends up. Connect the external tools, extract the feature requests, dedupe them against what's already in Canny, and return them as drafts for review.
It was a big swing for a bootstrapped company, and parts of it would have to reorient around this. So we validated first, in lean rounds across our customer base, looking for enough signal rather than the perfect signal.
Many sources in, fewer posts out
The 90% bar
We began exploring Autopilot in late 2023, when newer AI models made reliable extraction practical. We gave ourselves 6 months to reach public launch: 4 to build an MVP, 2 to run a closed beta.
To keep scope tight, I was embedded with engineering throughout: daily syncs, shared specs, reviewing builds as they shipped.

Technical planning with the engineers
We couldn't overbuild the bet or pull engineering focus from the rest of the product. And the design had a harder job than usual. Most people didn't trust AI in this context yet, so the interface had to do the trust-building on its own.
The foundation was young, too. I'd spent the four or five months before Autopilot building Canny's first product-wide design system, mostly by standardizing patterns the product already had. Autopilot was its first real stress test, a new surface category full of decisions the system didn't have answers for yet.
Teams needed Autopilot to be right about 90% of the time before they'd let it near their operations. So we benchmarked every frontier model against our own backlog. We used Canny internally ourselves, which meant years of already-processed feedback as ground truth. We built the eval in-house, scored each model against the human calls, and reran it until the results held.
Zendesk ticket · sample 412 / 1,600
The model's calls are scored against decisions the team had made. Each pass replays a new ticket.
Benchmarking is token-heavy though, so we stayed deliberate about where compute went.
The pipeline that kept benchmarking affordable. Cheap models filter every ticket so the frontier model only reads what survives.
The inbox
I studied tools built for high-volume processing (Intercom, Zendesk, MailChimp). They all share one pattern. An inbox, a visual queue you work through top to bottom until you reach zero.
Reviewers here were deciding whether to trust each item, so every suggestion had to sit next to the original feedback it came from. Enterprise teams couldn't reorient their operations around signals they couldn't audit.
My first pass was a flat table: every item in one stream, tagged by type, quick actions at the right edge. Testing surfaced the problems fast. Mixing decision types forced constant context-switching, and at real density the type labels stopped registering.
Every decision type in one stream
- 1One flat streamEvery extraction lands in a single table.
- 2Type badgesPills tag each row merge, new post, or spam.
- 3Quick actionsApprove or dismiss without leaving the row.
- 4Spam beside signalJunk sat in the same stream as real requests.
Each iteration went after whatever was costing reviewers the most. V1 split the stream into queues by decision type so similar calls could be batched. In V2, provenance moved to the front, so every suggestion could be traced to the conversation it came from. V3 pulled the density back to one clear decision per row, and that became the surface the closed beta used daily.
The beta taught us the next move. Users were accepting 91% of Autopilot's suggestions, just doing it manually. So I placed an automation prompt at the top of the queue, letting teams hand over decisions they were already agreeing with. Then the first customer let it run their whole queue, and the surface we'd built for checking every suggestion became an audit log.
The beta also showed the ceiling. However clean the rows got, row-by-row scanning wore people down as backlogs grew. So I designed and spec'd V4 in full, a browsable list beside a focused detail canvas. But V3 was performing too well to justify rebuilding the core surface, and V4 stayed on the shelf for a later release.
Pick from the queue, check what it matched on, then merge, publish or mark as spam. Every decision is undoable.
Extraction quality depended on how much the model knew about each team's product. So teams got the Knowledge Hub, where they upload reference material and extractions stay grounded.
Knowledge Hub: product context for grounded extractions
The design system got paid back too. Nearly all of Autopilot's UI fed the roughly 40-component library I'd been building: the cards, the banners, the side navs, the pills, the metadata rows. The next AI surface wouldn't have to start from zero.
A few sample molecules: cycle components, swap variants, toggle atoms, click text or a primitive to edit
Typeform's test
Typeform had the clearest proof point of the closed beta. They handle about 7,500 tickets a month from Zendesk across support, sales, onboarding, and CS. At that volume, manual tagging was never going to scale.
In a side-by-side review of 1,600 tickets, Autopilot processed the set in 53 minutes, surfaced 109 feature requests, and hit 93% accuracy, about 30 points higher than manual review. Deduplication held up too, with a 98.3% acceptance rate.
At that pace, Autopilot clears a full 7,500-ticket month in about 4 hours.
What it became
Autopilot changed how teams used Canny. The queue became somewhere they started their day.
- We logged 80% more feature requests after launch.
- Autopilot launched as a paid, usage-based add-on and cleared $400K in ARR in its first year.
Autopilot kept growing after launch, too. It became the umbrella for Canny's AI suite (Feedback Discovery, Smart Replies, Comment Summaries), and by mid-2025 it was folded into every plan.
Canny now sells itself as an AI-powered customer feedback platform. What we scoped as a six-month experiment is the first thing on the homepage.

Celebrating Autopilot's launch in Perugia