<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
    <channel>
        <title><![CDATA[Blog | Ankit Solanki]]></title>
        <description><![CDATA[Blog | Ankit Solanki]]></description>
        <link>https://ankitsolanki.com/blog/</link>
        <generator>RSS for Node</generator>
        <lastBuildDate>Wed, 31 Dec 2025 14:41:40 GMT</lastBuildDate>
        <atom:link href="https://ankitsolanki.com/blog/feed" rel="self" type="application/rss+xml"/>
        <item>
            <title><![CDATA[Predictions for 2026]]></title>
            <description><![CDATA[<h3>Coding</h3>
<ul>
<li>Coding agents will one-shot most problems put to them in sufficient
detail.
<ul>
<li><a href="https://zhengdongwang.com/2024/12/29/2024-letter.html#:~:text=The%20first%20awesome%20conclusion%20of%20the%20model%20does%20the%20eval%20is%20that%20we%20will%20achieve%20every%20evaluation%20we%20can%20state%2E">Anything you specify will get implemented</a>.</li>
<li>The next bottleneck will be deciding what to build, and the
difficulty of trading-off conflicting needs. A software paradox of
choice.</li>
</ul>
</li>
<li>Professional software developers will still produce the vast majority
of software in the world. Developers will continue to have jobs.</li>
<li>Developer productivity will only get a 10-15% boost <em>on average</em> by
most measures.</li>
<li>There will be outlier <strong>wizards</strong>: machine whisperers who produce
terrifying amount of software.</li>
<li>Companies start forcing people to adopt coding agents top-down, and it
doesn't go well. AI usage will start becoming a 'goal' and will get
<a href="https://en.wikipedia.org/wiki/Goodhart%27s_law">Goodhart-ed</a>.</li>
<li>There will be a movement to reject coding agents, reject LLMs and go
back to the good old days of writing code by hand.
<ul>
<li>It will fail to get traction in the workspace, where most software
is written for <a href="https://en.wikipedia.org/wiki/Instrumental_and_intrinsic_value">instrumental goals</a>.</li>
</ul>
</li>
</ul>
<h3>Knowledge work</h3>
<ul>
<li>Most white collar work will continue as-is. AI diffusion will remain
low.</li>
<li>Ham-fisted attempts to force AI use will lead to a growing backlash.</li>
<li>There will be multiple Claude Code equivalents for knowledge work, but
no clear winner.</li>
<li>By the end of the year, we'll reach a state where the case for using
agents for knowledge work simply can't be ignored.</li>
</ul>
<h3>Economy &#x26; Society</h3>
<ul>
<li>Anti-AI sentiment will be much higher in 2026.</li>
<li>The EU will seriously try to legislate AI use, to predictable
consequences.</li>
<li>There will be at least one major stock market crash and it will seem
like the bubble is bursting. This won't last though, and the bubble
will keep on bubbling.</li>
<li>Generative videos will be <em>really</em> popular even though no one really
thinks it's a great idea.</li>
</ul>]]></description>
            <link>https://ankitsolanki.com/blog/predictions-for-2026</link>
            <guid isPermaLink="false">https://ankitsolanki.com/blog/predictions-for-2026</guid>
            <dc:creator><![CDATA[Ankit Solanki]]></dc:creator>
            <pubDate>Wed, 31 Dec 2025 00:00:00 GMT</pubDate>
        </item>
        <item>
            <title><![CDATA[Reflections on 2025]]></title>
            <description><![CDATA[<p>Here's are a pair of pictures worth a thousand words:</p>
<p class="full-bleed"><img src="/media/github-contributions-2024.jpeg" alt="My Github contributions in 2024"></p>
<p class="full-bleed"><img src="/media/github-contributions-2025.jpeg" alt="My Github contributions in 2025"></p>
<p>This was the year of agents. We've been on a wild ride.</p>
<p>The period from June to December has been one of the most exhilarating
and productive times of my career, and I can't wait to share more about
the project I'm working on.</p>
<hr>
<p>Every group has its own vocabulary. Here's a set of words that makes
perfect sense to me:</p>
<ul>
<li>straight lines on a graph</li>
<li>the bitter lesson</li>
<li>jagged intelligence</li>
<li>scaling</li>
<li>giant inscrutable matrices</li>
<li>short timelines</li>
</ul>
<p>If you follow AI progress, you'll relate to some of these terms.</p>
<p>If you don't, we live in different worlds right now: but not for long.</p>]]></description>
            <link>https://ankitsolanki.com/blog/reflections-on-2025</link>
            <guid isPermaLink="false">https://ankitsolanki.com/blog/reflections-on-2025</guid>
            <dc:creator><![CDATA[Ankit Solanki]]></dc:creator>
            <pubDate>Wed, 31 Dec 2025 00:00:00 GMT</pubDate>
        </item>
        <item>
            <title><![CDATA[AI Art & Artistic Intent]]></title>
            <description><![CDATA[<p>You could look at any art in isolation, on its own merits; or in a
broader context, where you try to understand the <em>intent</em> of the
original artist, understand art as a product of its time and <em>interpret</em>
it in context of the broader artistic styles of the times.</p>
<p>I used to think that art should be completely separated from the artist.
The context shouldn't matter, the original author's intentions shouldn't
matter. The only thing that matters is how you react to an artwork.</p>
<p>A lot of modern art feels artists talking to each other and continuing a
conversation I'm not a part of. I believed that I could learn to
appreciate this form of art if I put in the effort, but I never had the
inclination to make the effort.</p>
<p>AI-generated art is somewhat changing my mind on this.</p>
<p>Sora 2, MidJourney, Veo 2: it seems like every day there's a new state
of the art method to generate images, sound or video. A lot of this is
slop, but a lot of this is also indistinguishable from human art. Often
it's better on a pure technical level.</p>
<p>My reaction to AI art is indifference for the most part. But not on its
own merits: I often come across a piece of writing or an image that I
think is cool on first glance: but on learning that it's AI-generated I
lose interest.</p>
<p>This has taught me a bit about myself: maybe artistic intent does matter
to me?</p>
<hr>
<p>I am an AI optimist: I think that the LLMs are miracles that we already
take for granted. I use AI daily, and I'm building-with &#x26; building-for
AI daily.</p>
<p>AI art sometimes takes my breath away. But I don't feel as connected to
it as a human artwork. For example, if I'm reading some fiction, I would
much rather read something written by a human even if AI could do a
better job.</p>
<p>I'm not sure why I feel this way, yet.</p>
<hr>
<p>Possible reasons for this?</p>
<ul>
<li>
<p>Maybe I don't like low-effort art? This implies that I don't value the
output itself, I also value all the inputs that go into creating
something.</p>
<p>Does this mean artistic intent does matter to me, after all? Is it the
effort that goes into creating something, or the amount of reps you
put in becoming good at what you do?</p>
</li>
<li>
<p>Maybe it's a defence mechanism? I know that the volume of AI generated
art will keep increasing exponentially, and you can only feel wonder
so many times a day.</p>
<p>Maybe avoiding AI art is a way to avoid becoming numb?
<a href="https://en.wikipedia.org/wiki/Wirehead_(science_fiction)">Wireheading</a> seems like a bad end after all.</p>
</li>
<li>
<p>Maybe we care about the stories behind the art, the process of making
something, instead of just the output?</p>
<p>These stories don't necessarily need to be real, they just need to be
shared. Narratives <em>are</em> powerful. And there's no <em>why</em> behind AI art.</p>
</li>
</ul>
<hr>
<p>For now, I have become more thoughtful about what I consume, and I'm
re-thinking some of my priors.</p>
<p>I dismissed modern art because I thought that context doesn't matter,
but now I'm realising that it does matter to me. The stories we tell
ourselves <em>do</em> matter.</p>]]></description>
            <link>https://ankitsolanki.com/blog/ai-art-and-artistic-intent</link>
            <guid isPermaLink="false">https://ankitsolanki.com/blog/ai-art-and-artistic-intent</guid>
            <dc:creator><![CDATA[Ankit Solanki]]></dc:creator>
            <pubDate>Tue, 14 Oct 2025 00:00:00 GMT</pubDate>
        </item>
        <item>
            <title><![CDATA[MCP Integration Patterns]]></title>
            <description><![CDATA[<div class="callout">
Note: This post was originally published on the <a href="https://cleartax.in/ai/">ClearTax AI blog</a>
</div>
<p>We're building an agent system and I have been evaluating adding support for MCP.</p>
<p>We <em>really</em> care about our tool design to an unreasonable degree. I have personally spent days thinking about the right design for some of our tools.</p>
<p>I couldn't help but over-think exactly how MCP should be exposed to our agents. At a high level, I felt that there could be two integration patterns — I would name them the <em>'meta tool pattern'</em> and <em>'materialised tools pattern'</em>.</p>
<h2>Meta Tool Pattern</h2>
<p class="full-bleed"><img src="/media/mcp-meta-tool-pattern.png" alt="Meta Tool Pattern"></p>
<p>The meta tool pattern: the agent sees a few meta tools like 'list available MCP servers', 'describe MCP server', 'invoke MCP tool'.</p>
<p>The agent can decide to explore the capabilities of an MCP server and invoke its tools, when necessary. This design is flexible and efficient, but also requires a more capable agent:</p>
<ul>
<li>This pattern uses less context by default. Tool definitions are lazily loaded only when required.</li>
<li>The 'discovery' calls for individual MCP servers need to happen only once in a given chat thread.</li>
<li>There's no guarantee that the agent will decide to explore the installed MCP servers though.</li>
</ul>
<h2>Materialised Tools Pattern</h2>
<p class="full-bleed"><img src="/media/mcp-materialised-tools-pattern.png" alt="Materialised Tools Pattern"></p>
<p>Here, whenever a MCP server is enabled – the system will automatically discover all available tools in the specific MCP server and eagerly copy them to the list of tools available to the agent. The agent sees all of these tools by default. Calling an MCP tool is just like calling any other tool.</p>
<ul>
<li>This pattern makes tools really explicit to the agent.</li>
<li>This comes at the cost of using additional context, even when it's not necessary.</li>
<li>If multiple MCPs are installed, you may fast run out of context space.</li>
</ul>
<h2>Evaluating these patterns</h2>
<p>If you were designing an agentic platform, which option would you choose? I did a survey of some existing open source systems and here's what I found:</p>
<ul>
<li><a href="https://github.com/openai/codex">codex-cli</a> also uses the materialised tools pattern</li>
<li><a href="https://github.com/sst/opencode">opencode</a> uses the materialised tools pattern</li>
<li><a href="https://github.com/cline/cline">Cline</a> uses a hybrid
<ul>
<li>All tools exposed by enabled MCP servers are copied to the system prompt</li>
<li>A single <code>use_mcp_tool</code> tool is used to invoke them</li>
</ul>
</li>
<li><a href="https://github.com/FoundationAgents/OpenManus">OpenManus</a> also uses the materialised tools pattern.</li>
</ul>
<p>This was a surprise. I'm not sure why the existing implementations are so heavily skewed towards materialised tools!</p>
<h2>Our Decision</h2>
<p>At this point I'm inclined to go with the <strong>meta tool pattern</strong> — it seems to make sense the system we're building.</p>
<p>What I have noticed is that:</p>
<ul>
<li>Most MCP servers don't have a great agent interface.</li>
<li>They expose too many fine grained tools with overlapping responsibilities.</li>
<li>They are really wasteful of context tokens.</li>
</ul>
<p>Most importantly: the meta tool pattern just intuitively feels like the <em>right</em> solution to me.</p>
<p>There are 100s of small details like this that go into building great products – and I have a feeling that these details are actually what differentiates your product when everyone is building on the same foundation models.</p>]]></description>
            <link>https://ankitsolanki.com/blog/mcp-integration-patterns</link>
            <guid isPermaLink="false">https://ankitsolanki.com/blog/mcp-integration-patterns</guid>
            <dc:creator><![CDATA[Ankit Solanki]]></dc:creator>
            <pubDate>Fri, 26 Sep 2025 00:00:00 GMT</pubDate>
        </item>
        <item>
            <title><![CDATA[The Mythical Agent Month]]></title>
            <description><![CDATA[<div class="callout">
Note: This post was originally published on the <a href="https://cleartax.in/ai/">ClearTax AI blog</a>
</div>
<p>The <a href="https://en.wikipedia.org/wiki/The_Mythical_Man-Month">Mythical Man Month</a> famously had the observation that adding
manpower to a project that's behind schedule will often delay it even
further. Additionally, in <a href="https://worrydream.com/refs/Brooks_1986_-_No_Silver_Bullet.pdf">'No Silver Bullet'</a> Fred Brooks further
states that:</p>
<blockquote>
<p>There is no single development, in either technology or management
technique, which by itself promises even one order-of-magnitude
improvement within a decade in productivity, in reliability, in
simplicity.</p>
</blockquote>
<p>How does this change with the advent of AI coding agents? Can coding
agents give us the mythical 10x speedup?</p>
<p>My thoughts below. I've divided this blog post into sections which argue
both for and against transformative change.</p>
<p>My position: coding agents are a fundamental shift in how we'll build
software over the next few decades, and we're overestimating the impact
in the short term while underestimating the impact in the long term.</p>
<h2>Headwinds</h2>
<h3>Vibe Coding vs Writing Production Software</h3>
<p>We should differentiate between 'vibe coding' as a way to experiment /
prototype, and using AI to write production software. <a href="https://simonwillison.net/2025/Mar/19/vibe-coding/">Simon
Willison</a> has an excellent post on this.</p>
<p>For production quality software, all basic tenets of software
engineering apply. Code reviews, tests, architecture designs, software
design reviews, etc.</p>
<p>Most of a senior engineer's time is often spent in these activities, not
just writing code.</p>
<p>Coding agents help a great deal here, but you won't see the same speedup
as pure vibe coding — where you build software <em>without reviewing the
code your agent writes</em></p>
<p>At least for now: coding agents aren't good enough to fully autonomously
build features and ship them to production without human review.</p>
<h3>Essential Complexity vs Incidental Complexity</h3>
<p>As <a href="https://worrydream.com/refs/Brooks_1986_-_No_Silver_Bullet.pdf">Fred Brooks pointed out</a> in the above essay, software development
consists of both essential complexity and incidental complexity.</p>
<p>Incidental complexity could be things like figuring out how to write a
Dockerfile, or learning how a specific library works, or dealing with
framework specific issues. Coding agents can be a huge help here.</p>
<p>Essential complexity is the core problem you're trying to solve. Coding
agents can definitely help here, but you still need to pay close
attention — humans will remain the bottleneck here.</p>
<p><a href="https://en.wikipedia.org/wiki/Amdahl%27s_law">Amdahl's Law</a> basically gives a ceiling for the performance gain
that automation / parallelisation gives for any given task. You are only
as fast as your bottleneck.</p>
<h3>Decision Fatigue and Time Compression</h3>
<p>Faster coding actually compresses timeframes and lets you focus on the
hard decisions, on the essential complexity. Coding agents let you focus
on substance of your problem.</p>
<p>But human capacity for deep thought is limited!</p>
<p>So now, your day to day working with AI coding tools is going to be a
series of hard-decisions that you need to think deeply upon, decisions
that require high amount of mental effort.</p>
<p>Decision fatigue is real. If you have to make a weeks' worth of hard
decisions in a day, your decision quality will suffer.</p>
<p>AI coding will exhaust you if you're not careful. Human beings need to
be able to step back and think about problems. We need to go for walks,
ruminate on ideas and just wander through a problem space.</p>
<h3>Effective Communication &#x26; User Skill</h3>
<p>AI agents need engineers to be effective communicators, and this is a
problem. Most engineers aren't the best communicators. Every great coder
isn't automatically great at delegation.</p>
<p>Effective communication is a skill. Writing clearly is a skill. And
using coding agents effectively is a skill.</p>
<p>For example, here are two recent articles that go in-depth about the
craft of using AI agents to code:</p>
<ul>
<li><a href="https://ampcode.com/how-i-use-amp">How I use Amp</a></li>
<li><a href="https://blog.nilenso.com/blog/2025/05/29/ai-assisted-coding/">AI-assisted coding for teams that can't get away with vibes</a></li>
</ul>
<p>Craftsmanship takes time. Skills take time to build. People who are
great engineers today won't automatically be great at using coding
agents. Getting better will require deliberate practice, and approaching
this problem with a beginner's mindset.</p>
<h3>Headwinds Summary</h3>
<p>Given that:</p>
<ul>
<li>Production software (currently) requires human oversight</li>
<li>Essential complexity remains</li>
<li>It will take time to learn how to use the coding agents effectively</li>
</ul>
<p>Is an immediate 10x improvement in velocity possible? It seems there is
truly no silver bullet.</p>
<h2>Tailwinds</h2>
<h3>AI Scaling will continue</h3>
<p>Agents will keep getting better. Underlying models will keep getting
better. We've learned to not bet against <a href="https://cleartax.in/ai/posts/will-scaling-continue">scaling</a>.</p>
<p>According to one recent viral benchmark, <a href="https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/">the length of tasks that AI can
do uninterrupted</a> is doubling every 7 months.</p>
<p>From my personal experience, I know that each recent big model release
(eg: Sonnet 3.5, Sonnet 3.7, Sonnet 4.0) has made building agents
easier. The LLMs are getting better at following instructions, at using
tools, at planning, and just at showing <em>agency</em>.</p>
<p>It's hard to predict the future, but it's definitely possible that soon,
a large majority of code written won't need human review and oversight.</p>
<h3>Most work isn't 'Deep Work'</h3>
<p>While my arguments above hold (essential complexity remains, decision
fatigue is real) — let's be real and acknowledge the fact that most of
us don't do deep work 100% of the time.</p>
<p>A lot of time goes into glue work, into getting various subsystems to
behave, dealing with broken tools, etc.</p>
<p>AI can be a huge accelerator for these types of work. It's possible that
this itself can be a 10x improvement for many organisations!</p>
<h3>Quantity has its own Quality</h3>
<p>If you're working in deep tech, if you're building something complex —
AI coding allows you to try more approaches. You can build quick and
dirty throwaway prototypes, and validate more ideas.</p>
<p>Quantity has a quality of its own. If you can do more iterations, you
can get to better decisions. If you can actually build multiple
candidate systems, you can make more informed choices — architecture
/ design decisions become easier with data.</p>
<p>I have personally seen this pay off: while building a zero to one
product, I have been able to do many parallel experiments and actually
test 10s of ideas before deciding upon a plan. This has enabled me to
make bold system design bets with high confidence.</p>
<h3>Ambition &#x26; Moonshots</h3>
<p>AI agents allow you to <a href="https://cleartax.in/ai/posts/be-more-ambitious">be more ambitious</a>. On the margin, with
higher productivity it's possible to devote more time and resources
towards building something <em>better</em> than you would have previously.</p>
<p>You can think of a productivity gain as either:</p>
<ul>
<li>Work 10x faster</li>
<li>Build something 10x better</li>
</ul>
<p>Either option is fine! In fact, it may be the case that building
something 10x better is actually more impactful.</p>
<p>I suspect one of the impacts of ubiquitous coding agents will be the
rising baseline quality of software!</p>
<h3>Tailwinds Summary</h3>
<p>If you consider the facts that:</p>
<ul>
<li>LLMs will continue to improve</li>
<li>We'll all get better at using AI agents</li>
<li>We'll be able to automate away low impact work</li>
<li>We'll be able to try many more iterations</li>
<li>We'll be able to build more impactful, more meaningful software</li>
</ul>
<p>How can you doubt the impact that coding agents will have?</p>
<h2>Conclusions</h2>
<p>I've argued both sides of this. My position is:</p>
<ul>
<li>We're radically underestimating coding agents</li>
<li>Most of us are not ready to adopt agents at scale</li>
</ul>
<p>I think impact and adoption will not be uniform. People will have
different lived experiences with AI tools, with some dismissing AI
coding as a fad, and some enthusiastically thinking of these agents as a
panacea for all their problems.</p>
<p>I think coding agents will have a huge impact that is overrated in the
short term, but underrated in the long term.</p>
<p>And I think that today, to get the most out of current generation
agents, you have to really dive deep and uncover their limits yourself.</p>]]></description>
            <link>https://ankitsolanki.com/blog/the-mythical-agent-month</link>
            <guid isPermaLink="false">https://ankitsolanki.com/blog/the-mythical-agent-month</guid>
            <dc:creator><![CDATA[Ankit Solanki]]></dc:creator>
            <pubDate>Fri, 27 Jun 2025 00:00:00 GMT</pubDate>
        </item>
        <item>
            <title><![CDATA[General Purpose vs Specialised AI Agents]]></title>
            <description><![CDATA[<div class="callout">
Note: This post was originally published on the <a href="https://cleartax.in/ai/">ClearTax AI blog</a>
</div>
<p>Should AI agents be general purpose or should they be built for a specific task?</p>
<p>I would say <strong>yes to both</strong>.</p>
<h3>Special Purpose Agents</h3>
<p>A special purpose agent is an agent built to do one (or a few things) very precisely. Examples:</p>
<ul>
<li>RAG on a knowledge base and answer questions <em>only from the knowledge base</em>.</li>
<li>Convert images to very specific structured data format (eg: convert image to a contact vCard).</li>
<li>Convert text to SQL on a specific table, with specific guardrails (eg: always apply page size limits).</li>
</ul>
<p>These agents are built to do a specific job.</p>
<h3>General Purpose Agents</h3>
<p>A general purpose agent on the other hand would be an agent that has <em>emergent behaviour</em>, that can use its tools to do something it wasn't explicitly programmed for.</p>
<p>This is best demonstrated by an example. I recently coded up a toy agent that runs on the command line that had the following tools available:</p>
<ul>
<li>List directory</li>
<li>Read file</li>
<li>Query file (via <a href="https://duckdb.org/">duckdb</a>)</li>
<li>Read PDF</li>
</ul>
<p>This agent was designed to have generic tools, and it was allowed to do multiple tool calls if needed. It wasn't given a specific goal.</p>
<p>The lack of specificity actually made the agent more useful! In the last few days, I have used it to do:</p>
<ul>
<li>K8S cost optimisation — given a CSV containing some kubernetes utilisation data, this agent helped me find low hanging cost optimisation options</li>
<li>Data entry for my own personal finance needs — given some account statements, the agent was able to help me digitise them in a format I use for tracking my expenses</li>
<li>Do Q&#x26;A on my meeting notes</li>
</ul>
<p>Interestingly, the agent showed a lot more <em>agency</em> than I was expecting. For example, it would often execute multiple tool calls in response to a general 'hello' message!</p>
<p class="full-bleed"><img src="/media/agent-hello.png" alt="Response to a &#x27;hello&#x27; message"></p>
<p>All the recent wow moments I have had with AI are usually with general purpose agents. They can end up giving you unexpectedly rich experiences.</p>
<h3>Use Cases &#x26; Trade-offs</h3>
<p>If you're building an AI enabled product, both special purpose agents and general purpose agents have their place.</p>
<ul>
<li>
<p>You sometimes want determinism (or close to determinism) and repeatability.</p>
<ul>
<li>For example: if you're building a support desk, you may want to always categorise tickets in a certain way.</li>
<li>In such scenarios, special purpose agents are really useful. You can treat them as "intelligence that's an API call away".</li>
</ul>
</li>
<li>
<p>Special purpose agents are limited in scope.</p>
</li>
<li>
<p>General purpose agents are where the power of AI agents becomes apparent.</p>
</li>
<li>
<p>General purpose agents are going to be more expensive to build.</p>
<ul>
<li>You may need smarter LLMs, you may need to spend tokens on reasoning.</li>
<li>General purpose agents could use up millions of tokens.</li>
</ul>
</li>
<li>
<p>Shipping general purpose agents requires you to have the right underlying infrastructure.</p>
<ul>
<li>The agent is as powerful as the tools it has access to. You need to build the right primitives to unshackle the AI model.</li>
<li>You need to build the right platform for the AI agents to leverage!
<ul>
<li>For example: if your system processes a lot of data, you may need to invest in robust parallel execution.</li>
<li>If the agent could 'ask any question' of your data, you have to think about indexing and database performance.</li>
</ul>
</li>
</ul>
</li>
<li>
<p>The ideal end state of a general purpose agent is a <a href="https://arxiv.org/abs/2402.01030">coding agent</a> — an agent that can write code for a specific task and then execute it.</p>
<ul>
<li>This could lead to potential security issues, and you have to look out for other types of abuse.</li>
<li>You might need to invest in primitives like sandboxing here.</li>
</ul>
</li>
</ul>
<p>I see this as a continuum — the more general an agent, the more powerful it is, but you need to invest proportionally into building the right safeguards.</p>
<p>Most use cases may not need a general purpose agent. You can probably start off building an AI-enabled product by just focusing on special purpose agents.</p>
<p>General purpose agents are also the most valuable ones though. And building general purpose agents will require you to invest proportionally into your platform's core primitives.</p>]]></description>
            <link>https://ankitsolanki.com/blog/general-agents-vs-specialised-agents</link>
            <guid isPermaLink="false">https://ankitsolanki.com/blog/general-agents-vs-specialised-agents</guid>
            <dc:creator><![CDATA[Ankit Solanki]]></dc:creator>
            <pubDate>Tue, 10 Jun 2025 00:00:00 GMT</pubDate>
        </item>
        <item>
            <title><![CDATA[Be More Ambitious]]></title>
            <description><![CDATA[<div class="callout">
Note: This post was originally published on the <a href="https://cleartax.in/ai/">ClearTax AI blog</a>
</div>
<p>Building a OCR / Document AI pipeline used to be hard work.</p>
<ul>
<li>Training OCR models</li>
<li>Building multi-step systems (eg: line segmentation, layout detection,
table detection)</li>
<li>Ensuring that these steps works nicely with each other</li>
<li>Adding heuristics and special cases</li>
<li>Doing a lot of testing to ensure your baseline is good enough</li>
</ul>
<p>Now: you could just give a PDF to Gemini Flash and get a 'good enough'
output in one API call. What used to take weeks / months can now be done
in hours.</p>
<p>With AI models becoming more capable — you can often get very far on
your problem statement <em>if you just try</em>. But if you are stuck in an old
mindset and think about some problems as difficult or time consuming,
you may not even attempt harder problems!</p>
<p>I don't think realisation has sunk-in yet. We still follow old patterns
of behaviour, we still mostly try to build the same things as in the
past.</p>
<p>Here's another recent example: I need to parse something reliably. The
'right way' to do this is to write a tokeniser / parser, but this is
time-consuming. In the past, I would have just depended on some basic
regexes and would have just got it working (and over time — fixed edge
cases as they came up).</p>
<p>This time though: I decided to do it in the 'right way'. I worked with
an AI coding agent and got a tokeniser and recursive descent parser
working in around an hour. The AI wasn't perfect — I definitely needed
to know the theory and know when to give the right inputs. Even though
this wasn't automatic, I ended up up with at least a 10x productivity
boost.</p>
<p>The hard lesson for me personally was:</p>
<ul>
<li>I needed to <em>let go</em> of my pre-conceived notions that a task can take
days / weeks.</li>
<li>I had to <em>decide</em> to build something ambitious.</li>
</ul>
<p>I need to keep reminding myself of the fact that <strong>hard things are now
easy</strong>, that the impossible may be actually possible.</p>
<p>I think this is problem where younger folks have an edge over more
experience people: with experience comes caution, but now is the time to
let go of caution and just build.</p>
<p>To everyone reading:</p>
<ul>
<li>I suggest partnering with one of the current frontier models
<ul>
<li>Use <a href="https://www.cursor.com/">Cursor</a>, or <a href="https://docs.anthropic.com/en/docs/claude-code/overview">Claude Code</a>, or <a href="https://windsurf.com/">Windsurf</a>, or <a href="https://v0.dev/">v0</a>, or <a href="https://lovable.dev/">Lovable</a>, or
anything else you want)</li>
</ul>
</li>
<li>Try to build something ambitious today!</li>
</ul>]]></description>
            <link>https://ankitsolanki.com/blog/be-more-ambitious</link>
            <guid isPermaLink="false">https://ankitsolanki.com/blog/be-more-ambitious</guid>
            <dc:creator><![CDATA[Ankit Solanki]]></dc:creator>
            <pubDate>Wed, 14 May 2025 00:00:00 GMT</pubDate>
        </item>
        <item>
            <title><![CDATA[AI Agents could be true 'User Agents']]></title>
            <description><![CDATA[<div class="callout">
Note: This post was originally published on the <a href="https://cleartax.in/ai/">ClearTax AI blog</a>
</div>
<p>In HTTP, browsers are called ‘<a href="https://www.w3.org/WAI/UA/work/wiki/Definition_of_User_Agent">user agents</a>’ — they were meant to act on
behalf of users. This initial positioning has sort-of carried forward
even now:</p>
<ul>
<li>
<p>Browsers are one of the few pieces of software left that allow plugins
&#x26; addons, that allow customisability.</p>
</li>
<li>
<p>Browsers let users <a href="https://support.mozilla.org/en-US/kb/change-fonts-and-colors-websites-use">override choices</a>. You can set your own min
font size, overriding any site’s CSS. You can use reader modes.</p>
</li>
<li>
<p>Browsers let users inspect applications via devtools, script
applications by allowing users to run code on a developer console.</p>
</li>
</ul>
<p>Browsers are one of the last vestiges of an old-school way of building
software: software that is extensible, that users can control and change
to better suit their needs.</p>
<p>Most current software we use today is the opposite.</p>
<ul>
<li>
<p>Most software (whether SaaS, mobile apps or desktop apps) is not
programmable.</p>
</li>
<li>
<p>Decisions are taken by product builders, not users. If a product owner
hasn’t added a feature that you want, you can’t go and add it
yourself!</p>
</li>
<li>
<p>Most software doesn’t empower users, it enforces rigid workflows.</p>
</li>
<li>
<p>We have taken away most customisability. Each option in a SaaS product
is expensive to support, and it’s rational as product builders to
reduce surface area and remove options that 99.9% of your users don’t
use.</p>
</li>
</ul>
<p>All of this happens because today’s software users aren’t programmers:
they don’t know the internals of how software, operating systems work.
Most software is built to be used by the mass market and as a result it
encodes a set of assumptions about how it's meant to be used.</p>
<p><strong>AI agents could reverse this trend.</strong></p>
<p>I can see a world where users have personal AI agents running on their
behalf. It wouldn’t matter if a product is scriptable or not, or if it
has an API or not — users could just ask the agent to automate a task on
their behalf! The agent may use the API, or it may just resort to
spinning up a browser and clicking on pixels.</p>
<p>What happens when most users are able to finally take control of the
tools that they use daily? I think it's finally time for general purpose
compute to be actually accessible by the masses.</p>
<p>This fills me with optimism; similar to the heady days of Web 2.0: when
everyone wasn’t so jaded about empowering users, exposing APIs was the
norm, mashups were possible and products like <a href="https://en.wikipedia.org/wiki/Yahoo_Pipes">Yahoo Pipes</a> were
built.</p>
<p>This is a perspective to keep in mind as we build AI-first products —
let's enable our users to do great things!</p>]]></description>
            <link>https://ankitsolanki.com/blog/ai-agents-user-agents</link>
            <guid isPermaLink="false">https://ankitsolanki.com/blog/ai-agents-user-agents</guid>
            <dc:creator><![CDATA[Ankit Solanki]]></dc:creator>
            <pubDate>Tue, 29 Apr 2025 00:00:00 GMT</pubDate>
        </item>
        <item>
            <title><![CDATA[Enabling AI Adoption]]></title>
            <description><![CDATA[<div class="callout">
Note: This post was originally published on the <a href="https://cleartax.in/ai/">ClearTax AI blog</a>
</div>
<blockquote>
<p>TL;DR: I'm sharing some ways in which we at Clear have democratised AI
adoption in the company; by giving people access to frontier models,
reducing dev friction, and actively fostering a learning culture around
AI.</p>
</blockquote>
<p>Here are some things we have tried at Clear to enable everyone in the
team to adopt AI tools.</p>
<h2>Launch an internal chat tool</h2>
<p>We launched an internal chat tool — similar to ChatGPT or Claude. This
is to enable everyone in the team to have un-metered access to the best
AI models — for example, we just added support for <a href="https://deepmind.google/technologies/gemini/pro/">Gemini 2.5 Pro</a>.</p>
<p>This internal tool works on top of APIs provided by OpenAI, Anthropic,
Google, etc. There are a lot of open source chat tools that take API
keys and work on top of them — we're using <a href="https://openwebui.com/">Open Web UI</a>.</p>
<p>Here's why you should do this:</p>
<ul>
<li>
<p>Free tiers of ChatGPT, Claude have very limited access to the frontier
models. Usually, free tiers are limited to the 'mini' models. If your
team hasn't used the most capable models, they don't know what LLMs
are capable of.</p>
</li>
<li>
<p>It's very expensive to pay for the pro or team tiers of each provider.
For example: ChatGPT team is $25 per member. APIs are usually cheaper,
especially since most users won't be chatting enough to cost you $25
in API credits.</p>
</li>
<li>
<p>Using APIs gives you flexibility to work with multiple providers in
parallel: it makes no sense to pay $25 for ChatGPT, $25 for Claude,
$25 for Gemini individually.</p>
</li>
<li>
<p>Free tiers are also risky from a security and privacy point. When you
use the free versions, most providers reserve the right to train on
your chats and use them to improve the models. In a corporate
scenario: having employees use personal accounts to do work is very
risky from a compliance perspective.</p>
</li>
</ul>
<h2>Make it easy to build AI projects internally</h2>
<p>We built an internal proxy which can be used by our developers to access
providers like Anthropic, OpenAI, etc. Both directly and via AWS
Bedrock, Microsoft Azure.</p>
<p>This drastically reduces friction when building an AI-powered use case:
developers can quickly do POCs and experiments without asking someone to
provision a specific API key on their behalf.</p>
<p>We learned this the hard way: during one hackathon that we ran, I
personally had to spend a lot of time just making sure that every team
had the required access!</p>
<p>There are many open source solutions for this. We use something custom,
but right now it looks like <a href="https://docs.litellm.ai/docs/">LiteLLM</a> is a great option.</p>
<p>Enabling our team to quickly build AI-powered applications, without
asking for permission, has proven to be a huge win. Sometimes, the best
ideas come bottom up — for example, one team used this capability to
enable automatic code reviews on Github powered by <a href="https://aider.chat/">aider</a>.</p>
<h2>Pedagogy and Evangelism</h2>
<p>There is also a need to do deliberate culture shaping — making sure you
are bringing everyone in the team along with you on the journey.</p>
<p>Examples of this include:</p>
<ul>
<li>
<p>Writing internal notes on AI — this blog actually started out as
internal writings. The field is very fast moving, keeping up with AI
progress is almost a full time job.</p>
<ul>
<li>
<p>There is a lot of value in having some team members synthesise
their learnings and sharing with your broader team.</p>
</li>
<li>
<p>There is also a lot of value in simplifying jargon and giving clear
explanations of what different AI concepts mean.</p>
</li>
</ul>
</li>
<li>
<p>Showcasing use cases, doing demos: people learn by association. If you
come across a novel way to use AI, it should be shared broadly in your
team! The best ideas will come from your team, find a way to enable
these ideas more broadly.</p>
</li>
</ul>
<p>The recent <a href="https://x.com/tobi/status/1909251946235437514">Shopify memo</a> is a great example of this in action.</p>
<hr>
<p>This is just the start. We are still learning about how to best adopt AI
in our daily lives. I'd be very curious to learn how everyone else is
approaching this problem.</p>
<p>Please reach out to me on email (ankit at clear dot in) or on
<a href="https://x.com/_anks">twitter</a>.</p>]]></description>
            <link>https://ankitsolanki.com/blog/enabling-ai-adoption</link>
            <guid isPermaLink="false">https://ankitsolanki.com/blog/enabling-ai-adoption</guid>
            <dc:creator><![CDATA[Ankit Solanki]]></dc:creator>
            <pubDate>Thu, 10 Apr 2025 00:00:00 GMT</pubDate>
        </item>
        <item>
            <title><![CDATA[Writing for AI]]></title>
            <description><![CDATA[<div class="callout">
Note: This post was originally published on the <a href="https://cleartax.in/ai/">ClearTax AI blog</a>
</div>
<p>Lately, it seems everyone uses LLMs to write. People are justifiably annoyed at the 'default' ChatGPT style of writing. Having ChatGPT write an email for you, or fix your grammar is a low value use case though.</p>
<p>I think there is a lot of value in doing the opposite: <strong>doing more writing yourself</strong>, so you can then feed it to an LLM. Writing not as an end, but to generate (quality) content for AI. Writing down things so that AI can help you (or others) down the line.</p>
<p>Interestingly: people like <a href="https://marginalrevolution.com/marginalrevolution/2025/01/should-you-be-writing-for-the-ais.html">Tyler Cowen</a> and <a href="https://gwern.net/llm-writing">gwern</a> go as far as to say that writing publicly for LLMs is functionally a way to gain immortality.</p>
<p>In a world where written knowledge is easily accessible by LLMs: there is a lot to gain when you pivot to a writing culture. Instead of making decisions face-to-face, switch to writing memos and jotting down your decisions explicitly. You can then use AI to reflect on your writing, give you deep feedback, get it to spot patterns and trends, and dig up the nuggets of insight that are often hard to find.</p>
<p>Here are some trends I've noticed recently:</p>
<ul>
<li>
<p>I have noticed people who are clued-in have started to take more notes than before.</p>
<ul>
<li>
<p>For example: apps like <a href="https://www.granola.ai/">Granola</a> let you record your meetings, take notes yourself, and then ask questions about your notes.</p>
</li>
<li>
<p>Even with Granola, the high value use cases come about when you jot down your own thoughts as well, not just depend on it to transcribe the meeting.</p>
</li>
</ul>
</li>
<li>
<p>There’s a trend to make API documentation, user guides etc all available as plain text in an <a href="https://llmstxt.org/">llms.txt</a> file. This documentation can then be easily fed to LLMs so that the LLM can help users write code.</p>
</li>
<li>
<p>Writing <a href="https://docs.anthropic.com/en/docs/build-with-claude/develop-tests">evals</a> is one of the highest value tasks you can do when building AI applications. People have been paid by AI labs to create high quality content and high quality evals for AI.</p>
<ul>
<li>Epoch AI for example <a href="https://www.reddit.com/r/math/comments/1h6rwls/im_developing_frontiermath_an_advanced_math/?rdt=45933">hired 70 mathematicians</a> to generate problems for their <a href="https://epoch.ai/frontiermath">frontier math eval</a>. Imagine being one of the top experts in your field, and the highest value use of your time isn’t doing research, but just  writing tests for frontier AI models!</li>
</ul>
</li>
</ul>
<p>Now, fit this pattern into your organisation. Writing tests, background material, context for AI shouldn't be a low value task; and you should put your best people on it.</p>
<p>Every organisation has a lot of <a href="https://en.wikipedia.org/wiki/Tacit_knowledge">tacit knowledge</a>, which is completely invisible to AI. It's not going to be possible to codify every little part of your business — but being deliberate about writing will help!</p>
<h2>Building tools to reflect on your writing</h2>
<p>For personal use, you don't need to be too fancy here. With LLMs having huge context lengths like Gemini, I can put every personal note I’ve ever written over the last decade into one LLM prompt.</p>
<p>As an organisation: you could start collating information across multiple internal systems (eg: PRDs, design documents, internal tickets, etc) and consolidating it in a central space, and adding AI tools on top. But before you do this, you need to pivot your org to a writing culture.</p>
<hr>
<p>Start writing things down — not just for yourself or your co-workers right now, but the AI agents you'll soon be working with. We have started doing this at Clear.</p>]]></description>
            <link>https://ankitsolanki.com/blog/writing-for-ai</link>
            <guid isPermaLink="false">https://ankitsolanki.com/blog/writing-for-ai</guid>
            <dc:creator><![CDATA[Ankit Solanki]]></dc:creator>
            <pubDate>Thu, 27 Mar 2025 00:00:00 GMT</pubDate>
        </item>
    </channel>
</rss>