Recently, we have seen our product development velocity pick up significantly. In the first two months of 2026 alone, we shipped five major features:
- Peers, peer fund benchmarking. LPs can tag comparable funds to create bidirectional benchmarking relationships and compare key metrics directly in their pipeline.
- GP Discovery, a curated database with 1,400+ GPs and 10,000+ funds, filterable by sector, geography, asset class, and quality scores.
- Similar GPs, matching GPs based on similarity, metadata, and fund characteristics to help LPs expand their sourcing pipelines.
- AI Search, AI-powered GP research. Natural language queries that perform real-time web research and return structured GP profiles with streaming narrative analysis.
- Vision, portfolio monitoring. An aggregated dashboard for cash flows and performance metrics (IRR, DPI, TVPI, NAV) with drill-down into individual fund-level returns.
One metric we see this in is our Average time to merge, which is a decent proxy for responsiveness and flow efficiency from opening a PR to getting review to merging. This was down 29% when looking at the last 90 days.
As a small team, the improvements in development speed most likely come down to two things: the release of Claude Opus 4.6 and GPT-5.3-Codex, and us investing time in our agentic development setup. This post focuses on the latter, how we have structured our workflow around AI agents to accelerate our product development.
How We Build With Agents

Our development workflow can high level be broken into three stages, Shaping, Development, and Test & Review, each with AI agents embedded into the process. Here is how it works in practice.
1. Shaping
Shaping is where work gets defined before any code is written. Inputs come from multiple streams: product strategy, customer calls, prospect conversations, and internal brainstorming, and most of it surfaces in Slack, which is where our team communicates day-to-day.
We use Linear as our source of truth for projects and issues. To bridge the gap between Slack conversations and Linear, we have two agents integrated directly into Slack.
Agents with context
The Linear bot lets us tag it in any Slack thread to automatically turn a discussion into a Linear issue or add it to an existing project. Context from customer calls, strategy discussions, or ad-hoc brainstorming gets captured as structured work items without anyone having to manually write tickets.
OpenClaw is an open-source framework for building internal bots. We implemented our own instance of it called Clawsten, named in honor of our senior advisor. It has deep context on our business, strategy, and product. After a client meeting, we can feed it a transcript and ask it to generate a project brief or issue, which then gets posted back to Slack and into Linear. It is particularly useful for translating unstructured meeting notes into well-scoped work.
Both agents live where the conversations already happen, so there is no context-switching. We add, update, and remove projects and issues in plain language directly from Slack threads. We also found that the agents are great at distilling long transcripts and unstructured meeting notes into well-scoped projects and issues. Claude and Codex can then pull from Linear projects and issues during development, and update them based on the work they do, so shaping is not a one-time handoff but a continuous loop.
2. Development
For development we use Claude Code and Codex, either through the Claude Code terminal or the Codex app, depending on individual preference. Both models have gotten remarkably capable since Claude Opus 4.6 and GPT-5.3-Codex, and we are regularly seeing production-grade code that can be deployed with confidence.
That said, model quality alone is not enough. What we have found is that the better the context you give these models, the better they perform. What is known as context engineering has become a real discipline for us. Three things have made the biggest difference.
Connecting development to shaping via MCP
We have connected Claude and Codex to Linear using MCP, so they can pull projects and issues directly from our project management system. This means that when an agent starts working on a task, it already has access to the well-scoped requirements that were shaped in the previous stage: the project description, acceptance criteria, and relevant context. This connection between shaping and development is essential. Without it, you would need to manually copy requirements into prompts every time.
Making the repo agent-ready
We have invested significant time in making our repository easy for agents to navigate. This includes structured CLAUDE.md and AGENTS.md files that specify our tech stack, testing conventions, branching strategy, security guardrails, and other fundamentals. The reason this matters is that agents start from scratch every time. Without these files, they would have to rediscover how the project is set up on every session. Having it clearly documented means they can get productive immediately.
One thing we are still iterating on is keeping these files high-level and using references to point agents to specific places in the codebase where they can learn more. It becomes more of a dictionary than a manual, enough to orient them without overloading the context window.
Commands and skills for repetitive tasks
For any development task we find ourselves doing more than once a day, we have built custom commands and skills. These standardize repetitive workflows so the output is consistent and we do not have to re-explain the same process each time. This has been a quiet but meaningful accelerator, it removes friction from the small tasks that add up over a day.
3. Test & Review
As our development speed increased with agents, testing quickly became a bottleneck. We were running unit tests, but end-to-end testing was still manual and started consuming a lot of time. So we focused on tooling, and two things made the biggest difference.
End-to-end testing
For all business-critical flows in the system, we now have Playwright tests. These spin up the full application and run through the actual interactions: clicking buttons, navigating workflows, testing the paths that matter most to our customers. This was a key enabler for shipping with confidence. Before Playwright, every code change meant someone had to manually verify that the important workflows still worked. Now that happens automatically, and it lets us move faster without second-guessing whether something broke.
Multi-model code reviews
Since code is cheap and easy to generate with agents, the review layer becomes even more important. You want to catch vulnerabilities and significant bugs early, before they get anywhere near production.
What we have found is that a multi-model approach works well here. Having one agent review its own code tends to miss things, it has the same blind spots that produced the code in the first place. So since we are using both Claude and Codex for development, we apply different agents as reviewers, including Greptile and Bugbot. A different model looking at the code brings a different perspective and catches issues the original author would not.
What's Next
Right now, our agents amplify our development time, they help us produce more, ship faster, and write better code. But they only work when we work. We are still steering them, guiding them through tasks, and that means when we log off for the day, development stops.
The next big step we are exploring is making agents work around the clock. We already have a backlog full of shaped projects and well-scoped issues. There is no reason agents cannot work through them overnight, so that when we come in the next morning, there are PRs ready to review, tests already passing, and code waiting to be merged.
That shift, from agents as amplifiers to agents as autonomous contributors, is going to be a step function in how fast we can ship. We are very inspired by OpenAI's Harness Engineering post and writings from Michael Truell, the CEO of Cursor, on this new era of software development.
More to come. We will keep you updated.