Wiring a Consensus-Based Agentic Stock Picker with Anthropic and Open Models

In my previous post Building an Agentic Stock Picker with Anthropic, I implemented an agentic stock picker built with Claude Opus 5. The application analyses NSE stocks under strict governance. The rule everything follows from is simple: the market-data connection and a running AI model are never live at the same time. Data is downloaded first and frozen on the laptop; only then do the agents run, with no route back to the data provider. In the agentic stock picker Sonnet writes the fundamental and technical claims besides acting as the auditor while Haiku handles the news.

Here the claim is a single checkable assertion about one company for e.g. a sentence citing figures from the frozen data (“MCX trades at a P/E of 55.00x against a sector P/E of 32.62x, so it is expensive on earnings”), paired with the tripwire that would prove it wrong. Since, the Claude runs cost money, I also wanted to explore Open Models.

As mentioned in my earlier post, I wanted to run the agentic stock picker on Open Models hosted entirely on my own machine, with no data sent anywhere and no cost per run. I tried Qwen, GLM, GPT-OSS, Mistral, Magistral, Gemma and Nemotron, and compared every one against Claude on the same set of stocks. No single open model matched Claude, so the post ends with a consensus of three open models from different labs: Google’s Gemma, Zhipu’s GLM and Alibaba’s Qwen.

Small models exposed weaknesses that a strong model had quietly papered over, and fixing them made the whole system better, Claude runs included.

Disclaimer:Note: This is not investment advice and has been done for purely research purposes.

The setup

  • Machine: a MacBook Pro with an M3 Max chip and 64 GB of memory.
  • Model server: Ollama, which runs open models locally and serves them on the laptop’s own network interface.
  • Every model analysed the same 15 shortlisted stocks from the same frozen snapshot. The results were compared with Claude’s on that snapshot. Claude was used as the reference. I wanted to see if a local model can reach similar conclusions.

Two numbers do most of the comparing:

  • Rank agreement asks whether the model puts the 15 stocks in the same order, best to worst, as Claude. 1.0 means an identical order; 0 means no relationship.
  • Score gap asks whether the scores themselves are close: the average difference from Claude’s, as a fraction of how spread out Claude’s scores are. Lower is better.

How the agents work (a quick recap)

Each AI agent (fundamental, technical and news) looks at one company and writes claims, such as “MCX is expensive on earnings: P/E 55.00 against a sector P/E 32.62”. Every claim must carry its own falsifier/tripwire, which if evaluated to true would result in the claim being discarded. We can think of the falsifier/tripwire of the claim as the negation of the claim. For that claim, the tripwire/negation is “if P/E is below the sector’s P/E”.

Python then checks every tripwire against the frozen data:

  • If the claim’s negation/tripwire is true, the claim is proven wrong and discarded.
  • If the tripwire could never be true, the claim can’t be tested, so it’s discarded too.
  • Only claims that could have been proven wrong, but weren’t, survive and count towards the score.

An auditor model then tries to knock down the survivors, and plain arithmetic turns what’s left into a score and a BUY, WATCH or AVOID.

Ten local models were tried on the same 15 stocks, and seven were set aside:

ModelLabRank agreement with ClaudeOutcome
Gemma 4 31BGoogle0.87Member, the strongest local model
Gemma 4 26BGoogle0.74Good and faster, but from the same family as the 31B
GLM-4.7-FlashZhipu0.71Member, fast (about 25 minutes)
Qwen 3.8 27BAlibaba0.57Member; weaker alone, but the average is improved by it
MagistralMistral0.33Dropped; the average was made worse by adding it
Mistral Small 3.2Mistral0.28Dropped; TITAN was called a BUY
Nemotron 3.5 LightningNVIDIA0.05Dropped; 12 of 15 stocks were rated BUY
gpt-oss 20B, Qwen 3 8B / 14BOpenAI, Alibaba(earlier data)Dropped; too optimistic, or superseded

Claude’s run (claude-sonnet-5 with claude-haiku-4-5, about $1.92 for 15 stocks) was used as the reference throughout.

1. Tripwires pinned to the stock’s own value

Qwen 8B kept setting a tripwire’s threshold at the stock’s current value, copied from the data. TITAN’s 30-day return was −4.112031%, and the model wrote:

“The price has underperformed … by over 4 percent.” Falsifier: if return_30d_pct > -4.112031

Since the return is exactly −4.112031, the tripwire (falsifier) can never go off, so the claim can never be proven wrong. The checker correctly discards it as untestable. But 29 of the 35 pinned tripwires were on claims against a stock, so the negative claims were hit far harder than the positive ones. In the end, only 24 of 72 negative claims survived, against 70 of 92 positive ones, and every stock’s score was pushed upward.

2. Bigger model

Qwen 14B works better and does not pin negation on exact values

Qwen 14B was overly optimistic and one-sided. It wrote 95 claims in favour of stocks and 48 against, and found none of the four stocks Claude marked AVOID. On the stock MCX it praised strong returns and growth, all true, but never mentioned a P/E of 53 against a sector average of 32. This was corrected by also checking on the downside risks. After this the claims swung from 95 for and 48 against to 38 for and 64 against, with 7 AVOIDs where Claude had 4.

3. Claims written backwards

Claim: “P/E is below the sector P/E” (cheap on earnings) Tripwire: wrong if pe <= sector_pe

Here the claim and the tripwire are in the same direction which is incorrect. The tripwire is actually the correct test for the opposite claim, “expensive”. Python evaluated it faithfully: is 70.7 ≤ 6.0? Since the tripwire did not fire a false “cheap” claim survived as a point in the stock’s favour. Three such claims helped push ACUTAAS to a BUY.

4. The signal table

The fix for claims written backwards was to stop asking the model to invent rules. A python based signal table is created that defines 26 indicators with 56 rules. Each indicator has a rule for owning the stock and a rule against, and each rule comes with its exact tripwire/negation.

IndicatorFORAGAINST
P/E vs sectorbelow the sector: “cheap on earnings”above the sector: “expensive on earnings”
ROCE vs sectorabove the sectorbelow the sector
Debt trendliabilities-to-equity fallingrising
RSI-14≤ 30 oversold; 30–45 in an uptrend (a pullback); 55–70 bullish≥ 70 overbought; 30–45 in a downtrend
MACDMACD line above its signal linesignal line above the MACD line

The signal table supplied a fixed, round threshold for every rule, so a stock’s exact reading can no longer become a tripwire. Automated tests prove that every tripwire is the exact opposite of its rule, including at the edges, where mistakes like > versus >= or AND versus OR hide. The table also carries 105 worked examples from real stocks, always written figures first, conclusion after: “GMRAIRPORT trades at a P/E of 128.31x against a sector P/E of 56.75x, so it is expensive on earnings.”

Sometimes the Local LLM auditors turned out to delete true claims. Some auditors refuted by copying the claim’s own, negation minus a clause. GLM-4.7-Flash’s auditor argued:

5. Auditor changes

“ANANDRATHI’s P/E (78.50x) is higher than the sector P/E (49.79x), contradicting the claim of being expensive.”

That argument restates and confirms the claim rather than contradicting it. I added two rules:

  • A challenge that merely restates the claim can’t remove it. This is checked logically, by testing whether the challenge’s condition follows from the claim’s.
  • A comparison Python has already computed (78.50 > 49.79) can only be challenged with other evidence.

These changes were replayed on stored runs and these rules now blocked 7 of GLM’s 9 wrong vetoes, and none of Claude’s.

6. Weighting
The local LLM models still rated the shortlist well above Claude. The reasons

Valuation counted no more than a price trend.

Every model gave almost every claim 0.85–0.95 confidence, whatever the evidence. A bigger model didn’t help: Qwen 27B was exactly as confident as Qwen 14B.

So confidence for rule-based claims is now computed from the data alone for e.g. 0.55 when a indicator is just past the rule’s threshold, up to 0.90 when it’s well beyond. A P/E of 25 against a sector P/E 24 gets 0.58, similarly a P/E of 53 against sector P/E 32 gets 0.90.

Now the weight is assigned based on the group to which the indicator belongs

GroupWeightCounts for the stockCounts against the stock
Valuation1.5P/E, EV/EBITDA or P/B below the sector’sP/E, EV/EBITDA or P/B above the sector’s
Debt (leverage)1.25Liabilities-to-equity fell over the year (not banks)Liabilities-to-equity rose over the year
Returns on capital1.0ROCE, ROE or ROA above the sector’s (ROCE not for banks)ROCE, ROE or ROA below the sector’s
Growth1.0Revenue or net profit up on the previous yearRevenue or net profit down on the previous year
Margins1.0Operating margin up over the yearOperating margin down over the year
Bank quality (banks only)1.0Net interest margin above the sector, net NPA below it, or CASA above itThe reverse of any of these
Momentum1.0RSI-14 ≤ 30 (oversold); 30–45 while above the 200-day average (a pullback in an uptrend); 55–70 (bullish). MACD above its signal lineRSI-14 ≥ 70 (overbought); 30–45 while below the 200-day average. MACD below its signal line
Risk1.0Within 10% of the 52-week high; volatility below 20%; ATR below 2% of priceMore than 20% below the 52-week high; volatility above 35%; ATR above 4%
Trend0.75Above its 200-, 50- or 20-day average; MACD above zeroBelow those averages; MACD below zero
Price returns0.5Up over 30 days, 90 days or a yearDown over those periods
Volume0.5A 5-day rise on rising volumeA rise on thinning volume, or a fall on rising volume

The Stock Picker approach

The whole pipeline, as it runs today, whichever model is plugged in.

Flowchart outlining stock selection and ranking process for Nifty 500, including filtering criteria and model evaluations, with a final consensus of stock recommendations.

Stage 1: Filters (pass or fail)

Each of the roughly 500 Nifty 500 stocks must pass every filter:

FilterSetting
Return on equity (ROE)at least 12%
Return on capital employed (ROCE)at least 12% (not applied to banks)
Revenue growthnot shrinking
P/Eat most 80
RSI-14between 30 and 80
Daily price range (ATR)at most 5% of the price
Data freshnessresults under 400 days old, prices under 5 days

Missing data counts as a fail. Around 190 stocks usually pass. Debt isn’t a filter here; it’s judged later, by its trend.

Stage 2: Ranking to a shortlist of 15

The survivors are ranked on 14 measures in five groups, each measure converted to a percentile against the whole market rather than tested against a fixed cut-off:

GroupWeightMeasures
Quality30%ROCE vs sector, ROE vs sector, margin trend
Value20%P/E, P/B and EV/EBITDA vs sector (lower is better)
Growth20%revenue growth, profit growth
Trend20%price vs 200-day average, MACD, 90-day return
Risk10%ATR, volatility, distance from 52-week high

A measure a company can’t have, such as ROCE for a bank, scores at the median rather than the bottom. The top 15 go forward.

Stage 3: The AI analysts

For each of the 15, Python works out which of the 56 rules hold.

Three analysts per stock. Each of the 15 shortlisted stocks is examined by three separate analysts. Each analyst is a separate call to the AI model, and each sees only its own kind of evidence:

AnalystLooks atRules from the signal table
Fundamentalvaluation, returns on capital, growth, margins, debt, bank measures13 indicators, 26 rules
Technicalprice trend, momentum, returns, risk, volume13 indicators, 30 rules
Newsrecent headlines about the companynone; its claims are written freely

Each analyst chooses up to six claims, including both sides whenever both exist, and writes the argument. The news analyst reads the headlines separately. Python checks every claim, and the auditor challenges the survivors under the rules above.

The 56 rules. The rules are kept in a Python file called the signal table. For each of 26 indicators, one rule is given for owning the stock (FOR) and one against it (AGAINST).

A few rows of the 26 rules used by the fundamental analyst a few are shown below

IndicatorFORAGAINST
P/E vs sectorcheap on earningsexpensive on earnings
EV/EBITDA vs sectorcheap on cash profitexpensive on cash profit
P/B vs sectorcheap on book valueexpensive on book value
ROCE vs sector (not banks)returns on capital above the sectorbelow the sector
ROE vs sectorreturns to shareholders above the sectorbelow the sector

A few of the 30 rules used by the technical analyst below

IndicatorFORAGAINST
RSI-14oversold (≤ 30); pullback in an uptrend (30–45, above the 200-day average); bullish (55–70)overbought (≥ 70); weak in a downtrend (30–45, below the 200-day average)
Price vs 200-day averagelong-term uptrendlong-term downtrend
Price vs 50-day averagemedium-term uptrendmedium-term downtrend
Price vs 20-day averageshort-term strengthshort-term weakness
MACD vs its signal lineMACD bullishMACD bearish

Stage 4: Score and recommendation

Each surviving claim contributes:

± dimension weight × group weight × confidence

Fundamental claims weigh 1.0, technical 0.8 and news 0.5; the group weights are in the table above.

For example, a stock clearly expensive on P/E (53 against 32, confidence 0.90) loses 1.0 × 1.5 × 0.90 = −1.35. Clearly strong ROCE earns 1.0 × 1.0 × 0.90 = +0.90. A clear 200-day uptrend earns 0.8 × 0.75 × 0.90 = +0.54. One clear valuation concern outweighs a strong return on capital.

The recommendation rules:

LabelRule
BUYscore ≥ 1.0, and positive claims in at least two dimensions, and no surviving claim against the business at confidence 0.6 or more
AVOIDscore ≤ −0.5
WATCHeverything else

A single clear negative about the business blocks a BUY, and most high-quality companies are expensive, so BUY is rare by design.

The results for the seven open models against Claude

ModelMakerRank vs ClaudeScore gapBUY / WATCH / AVOIDRun time
Claude (API, about $1.90 per run)Anthropic——1 / 9 / 514 min
Gemma 4 31BGoogle0.870.451 / 13 / 1~1.5 h
Gemma 4 26BGoogle0.740.461 / 12 / 2~35 min
GLM-4.7-FlashZhipu0.710.693 / 11 / 127 min
Qwen 3.8 27BAlibaba0.570.692 / 11 / 21 h 35 min
MagistralMistral0.330.872 / 12 / 11 h 3 min
Mistral Small 3.2Mistral0.280.762 / 13 / 01 h 18 min
Nemotron 3.5 LightningNVIDIA0.051.5012 / 3 / 020 min

gpt-oss 20B (OpenAI’s open-weight model) was tested on the earlier data only. It tended to be too optimistic, with rank agreement 0.31, and wasn’t carried it forward. Neither did the Qwen 3 8B and 14B models, which the table above already covers.

  • Google’s Gemma 4 31B agreed with Claude best by a clear margin, and wrote no invalid claims at all. It had been released only days before these runs.
  • Some models are simply optimistic. Nemotron rated 12 of the 15 a BUY, including TITAN, the stock Claude rated worst. gpt-oss and Mistral leaned the same way to a lesser degree. What set them apart was a tendency to choose supporting facts over warning signs, even though every fact they were given was true.
  • Speed and quality are not the same axis. The fastest models (Nemotron, gpt-oss) were among the least aligned, while GLM-4.7-Flash and Gemma 4 26B were both fast and reasonably close.
  • Every model agreed on one stock. GESHIP (Great Eastern Shipping) was a BUY for Claude and every open model tested on this data. It had a P/E of 5.7 against a sector 15, returns well above its peers, falling debt and an upward trend.
  • Several models agreed on the other end. TITAN was an AVOID for Claude, GLM, Qwen 27B and Gemma 4 26B. Its business is excellent (ROE 32%, revenue up 45%), but at 27.6 times book value against a sector 3.3, with rising debt and fading momentum, the valuation and debt weights pull it down.

A consensus of models

Because no single open model matches Claude, I tried averaging several. Models from different labs tend to make different mistakes:

CombinationRank vs ClaudeScore gap
GLM + Qwen 27B0.800.57
GLM + Qwen 27B + Gemma 4 31B0.820.52

The GLM and Qwen pair beat both of its members, and adding Gemma 4 31B improved it further. The trio still falls short of Gemma 4 31B on its own (0.87), which argues for giving Gemma a strong voice. So a three-lab consensus (Zhipu, Alibaba and Google) is taken to decide which stock gets a BUY or AVOID only when most models agree. The two Gemma models together scored higher still (0.92), but coming from the same family, they’re more likely to share blind spots.

Flowchart illustrating a consensus process for evaluating different AI models, including Gemma 4, GLM-4.7-Flash, Qwen 3.8, and Claude, showcasing the steps of scoring and labeling per stock and the final consensus calculation.

Here is what we get when we run each LLM independently

StockClaudeGLM-4.7-FlashQwen 3.8 27BGemma 4 26BGemma 4 31B
IKSWATCH (+4.95)WATCH (+3.34)WATCH (+4.65)WATCH (+2.85)WATCH (+4.65)
GESHIPBUY (+4.48)BUY (+4.66)BUY (+4.19)BUY (+5.00)BUY (+5.79)
ENGINERSINWATCH (+2.63)WATCH (+3.25)WATCH (+1.06)WATCH (+0.36)WATCH (+3.00)
LODHAWATCH (+1.84)BUY (+2.44)WATCH (+2.21)WATCH (+1.03)WATCH (+1.85)
HINDZINCWATCH (+1.83)WATCH (+1.59)WATCH (+3.75)WATCH (+1.67)WATCH (+2.50)
GABRIELWATCH (+1.56)WATCH (+2.96)WATCH (+0.28)WATCH (+0.66)WATCH (+1.56)
HBLENGINEWATCH (+0.74)WATCH (+2.49)WATCH (0.00)WATCH (+0.96)WATCH (+0.63)
MCXWATCH (+0.36)WATCH (+0.88)WATCH (+1.18)WATCH (−0.17)WATCH (+0.88)
LLOYDSMEWATCH (+0.30)WATCH (+0.49)WATCH (+1.76)WATCH (−0.13)WATCH (+1.57)
LUPINWATCH (−0.49)WATCH (+1.51)BUY (+2.90)WATCH (+0.07)WATCH (+1.27)
PREMIERENEAVOID (−0.55)WATCH (+1.94)WATCH (+0.98)WATCH (+1.87)WATCH (+0.73)
MUTHOOTFINAVOID (−0.91)BUY (+3.25)AVOID (−1.51)WATCH (+0.23)WATCH (−0.13)
KALYANKJILAVOID (−1.07)WATCH (+1.57)WATCH (+2.28)AVOID (−1.15)WATCH (+1.09)
GLENMARKAVOID (−1.39)WATCH (−0.01)WATCH (+0.30)WATCH (−0.10)AVOID (−0.89)
TITANAVOID (−2.28)AVOID (−0.92)AVOID (−0.75)AVOID (−1.85)WATCH (+0.40)
BUY / WATCH / AVOID1 / 9 / 53 / 11 / 12 / 11 / 21 / 12 / 21 / 13 / 1
Rank agreement with Claude—0.710.570.740.87
Score gap vs Claude—0.690.690.460.45

The number in brackets is each stock’s score. A stock is BUY at a score of 1.0 or more, but only if it is positive in at least two of the fundamental, technical and news areas and has no strong negative business claim. It is AVOID at −0.5 or below, and WATCH otherwise.

The Consensus strategy

The whole pipeline is run separately by each AI model, and a score and a BUY, WATCH or AVOID are produced for every stock . The finished results are then combined by regular Python.

  • The order of the stocks is set by a weighted average of the models’ scores.
  • The label is decided by a weighted vote on the models’ own labels. A stock is made a BUY only if BUY was given by models holding at least half the weight, and an AVOID if AVOID was given by models holding at least half.
  • A BUY is blocked by a heavyweight objection. If AVOID is given by any model holding 0.3 or more of the weight, the stock cannot be made a consensus BUY.
  • Everything else is labelled WATCH.
ModelWeight of its vote, without ClaudeWeight of its vote, with Claude
Claudenot in the vote0.40
Gemma 4 31B0.40.25
GLM-4.7-Flash0.30.175
Qwen 3.8 27B0.30.175
Total1.01.00
  1. Order: the weighted average of the scores

weighted score = Σ (model weight × model’s score)

With Claude, for GESHIP: 0.40 × 4.48 + 0.25 × 5.79 + 0.175 × 4.66 + 0.175 × 4.19 = +4.79.

2. . The BUY test

A stock is made a consensus BUY only when all three checks are passed:

  • No veto: AVOID was not voted by any single model whose own weight is 0.3 or more (rule 3).
  • Majority: the BUY weight is 0.5 or more.
  • Score: the weighted score is +1.0 or more.

3. The veto: a BUY is blocked by a heavyweight objection

If AVOID was voted by any one model whose weight is 0.3 or more, the stock cannot be made a BUY, however large the BUY weight.

Who can cast a veto is therefore set by the weights:

VersionWeight of 0.3 or moreCan veto a BUY on its own
Without ClaudeGemma (0.4), GLM (0.3), Qwen (0.3)all three
With ClaudeClaude (0.40) onlyClaude only; Gemma (0.25), GLM and Qwen (0.175) can vote against a BUY, but cannot block one alone

4. The AVOID test

A stock that is not a BUY is made an AVOID when either check is passed:

  1. Majority: the AVOID weight is 0.5 or more, or
  2. Score: the weighted score is −0.5 or below.

Unlike BUY, AVOID can also be reached through the score alone

5. Everything else is WATCH

This includes an exact even split, where the BUY weight and the AVOID weight are both 0.5. When the models are divided, “watch’ should be fine.

Using weights (without Claude)**

Gemma (0.4)GLM (0.3)Qwen (0.3)BUY weightAVOID weightVeto?ResultExample
BUYBUYBUY1.00noBUYGESHIP
BUYBUYWATCH0.4 + 0.3 = 0.70noBUYillustration
WATCHBUYWATCH0.30noWATCH: 0.3 is short of 0.5LODHA
BUYBUYAVOID0.70.3yes: Qwen’s 0.3WATCHillustration
WATCHAVOIDAVOID00.3 + 0.3 = 0.6—AVOIDTITAN
WATCHBUYAVOID0.30.3—WATCH: no majority either wayMUTHOOTFIN
AVOIDWATCHWATCH00.4—WATCH: 0.4 is short of 0.5, score −0.27GLENMARK
WATCHWATCHWATCH00—WATCH, despite a score of +4.26IKS

The results on 28 Sep data (with prices to 25 Sep 2026)

A) Without Claude (Gemma 0.4, GLM 0.3, Qwen 0.3)**

StockGemma 4 31BGLM-4.7Qwen 27BWeighted scoreConsensusClaude (not voting)
GESHIPBUY (+5.79)BUY (+4.66)BUY (+4.19)+4.97BUYBUY
IKSWATCH (+4.65)WATCH (+3.34)WATCH (+4.65)+4.26WATCHWATCH
HINDZINCWATCH (+2.50)WATCH (+1.59)WATCH (+3.75)+2.60WATCHWATCH
ENGINERSINWATCH (+3.00)WATCH (+3.25)WATCH (+1.06)+2.49WATCHWATCH
LODHAWATCH (+1.85)BUY (+2.44)WATCH (+2.21)+2.14WATCHWATCH
LUPINWATCH (+1.27)WATCH (+1.51)BUY (+2.90)+1.83WATCHWATCH
GABRIELWATCH (+1.56)WATCH (+2.96)WATCH (+0.28)+1.60WATCHWATCH
KALYANKJILWATCH (+1.09)WATCH (+1.57)WATCH (+2.28)+1.59WATCHAVOID
LLOYDSMEWATCH (+1.57)WATCH (+0.49)WATCH (+1.76)+1.30WATCHWATCH
PREMIERENEWATCH (+0.73)WATCH (+1.94)WATCH (+0.98)+1.17WATCHAVOID
HBLENGINEWATCH (+0.63)WATCH (+2.49)WATCH (0.00)+1.00WATCHWATCH
MCXWATCH (+0.88)WATCH (+0.88)WATCH (+1.18)+0.97WATCHWATCH
MUTHOOTFINWATCH (−0.13)BUY (+3.25)AVOID (−1.51)+0.47WATCHAVOID
GLENMARKAVOID (−0.89)WATCH (−0.01)WATCH (+0.30)−0.27WATCHAVOID
TITANWATCH (+0.40)AVOID (−0.92)AVOID (−0.75)−0.35AVOIDAVOID

B) With Claude (Claude 0.40, Gemma 0.25, GLM 0.175, Qwen 0.175)**

StockClaudeGemma 4 31BGLM-4.7Qwen 27BWeighted scoreConsensus
GESHIPBUY (+4.48)BUY (+5.79)BUY (+4.66)BUY (+4.19)+4.79BUY
IKSWATCH (+4.95)WATCH (+4.65)WATCH (+3.34)WATCH (+4.65)+4.54WATCH
ENGINERSINWATCH (+2.63)WATCH (+3.00)WATCH (+3.25)WATCH (+1.06)+2.56WATCH
HINDZINCWATCH (+1.83)WATCH (+2.50)WATCH (+1.59)WATCH (+3.75)+2.29WATCH
LODHAWATCH (+1.84)WATCH (+1.85)BUY (+2.44)WATCH (+2.21)+2.02WATCH
GABRIELWATCH (+1.56)WATCH (+1.56)WATCH (+2.96)WATCH (+0.28)+1.58WATCH
LLOYDSMEWATCH (+0.30)WATCH (+1.57)WATCH (+0.49)WATCH (+1.76)+0.91WATCH
LUPINWATCH (−0.49)WATCH (+1.27)WATCH (+1.51)BUY (+2.90)+0.90WATCH
HBLENGINEWATCH (+0.74)WATCH (+0.63)WATCH (+2.49)WATCH (0.00)+0.89WATCH
MCXWATCH (+0.36)WATCH (+0.88)WATCH (+0.88)WATCH (+1.18)+0.73WATCH
KALYANKJILAVOID (−1.07)WATCH (+1.09)WATCH (+1.57)WATCH (+2.28)+0.52WATCH
PREMIERENEAVOID (−0.55)WATCH (+0.73)WATCH (+1.94)WATCH (+0.98)+0.47WATCH
MUTHOOTFINAVOID (−0.91)WATCH (−0.13)BUY (+3.25)AVOID (−1.51)−0.09AVOID
GLENMARKAVOID (−1.39)AVOID (−0.89)WATCH (−0.01)WATCH (+0.30)−0.73AVOID
TITANAVOID (−2.28)WATCH (+0.40)AVOID (−0.92)AVOID (−0.75)−1.11AVOID

**Please note

Disclaimer:The stocks listed here should not be used as investment advice. This has been done for research purposes.

Some findings

  • The clear calls are agreed by both. GESHIP is rated BUY and TITAN AVOID in both versions.
  • More caution is shown by the version without Claude. Where the models disagree, WATCH is assigned.
  • The weak end is mostly sharpened by adding Claude.
  • Only 15 stocks on one date have been tested. The data is at least 2 weeks old
  • This has been done for research purposes alone and is not a stock recommendation

See also

  1. Building a Flight Simulator with Claude Code, Opus 5
  2. Introducing IPL AI Oracle: AI that speaks cricket!!!
  3. Introducing QCSimulator: A 5-qubit quantum computing simulator in R
  4. Experiments with deblurring using OpenCV
  5. Singularity

To see all posts click Index of Posts

Building an Agentic Stock Picker with Anthropic

For the past two years I have been hearing a lot about agentic flows. Yet all my work with AI-assisted coding has been deterministic task flows, for high-performance, scalable systems.

So I decided to try my hand at the agentic approach, and build an agentic stock picker. The idea was that the application would download data for stocks, get the fundamental and technical indicators along with the news items for each stock, and then have agents:

  • filter the stocks down on the fundamental and technical indicators
  • read the news items and try to identify why a stock price was dropping or growing
  • come up with a final list — stocks to buy, stocks to watch, and stocks to avoid

The approach was to use regular python to

  • download four years of prices, statements news, and compute every indicator, fundamental and technical namely ROE, ROCE, P/E, RSI, MACD, volatility and the rest
  • filter ~500 companies down to 15, on those numbers alone
  • checks every claim the agents make against the frozen data, and deletes the ones the data contradicts
  • scores whatever survives, and produces the final buy / watch / avoid list

and agents

  • argue about one company at a time with one agent for the business, one for the price chart, one for the news
  • and every claim has to come with the condition that would prove it wrong
  • then a second model attacks the claims that survived, and can only attack — there is no field in its answer for agreeing

Agentic Stock Picker

The dashboard of the Agentic Stock Picker is shown below
On the left we have

  • Sliders: the sliders to choose fundamental and technical ratios to filter on
  • Max candidates to recommend
  • Model: Offline heurestic/Anthropic/Qwen 14b( TBD)
  • Ceiling (USD):
  • Compile spend plan: This option will give the spend estimate and the hard stop
  • Approve and Run : If the amount is fine we can approve the run
  • Past runs: The results of past runs can be selected here
A screenshot of a financial analysis interface displaying stock data, including symbols, ratings, scores based on fundamental and technical analysis, and various financial metrics.

Design Strategy

Governance is the backbone of the implementation.

The agents run offline. When a model is running, there is no network. When the network is live, there is no model.

The network being live and a running model is never true

Text describing the first pillar titled 'Sealed window', highlighting network access during a specific phase.

The goal of the agent-based approach was to put the agents under a strict governance plan. In fact the whole design is built around governing them.

The agents never touch the market data provider at all. The download happens first, in ordinary Python with no model involved, over a fixed list of seven read-only URLs using a single read-only Upstox analytics token. By the time an agent starts, that connection is closed, the token has been dropped from the environment, and the process the agent runs in cannot even import the code that would talk to Upstox. This is least privilege enforced structurally rather than promised.

So the agents cannot reach the market, cannot trade, and cannot see a portfolio. There is no order, holdings or funds code path anywhere in the repository. While a model is running, exactly one address is reachable: the model provider itself.

The budget works the same way. The full cost of a run is computed and approved before the first call, and every call has to claim a one-use ticket against it. Going over budget is not something the system detects and recovers from. A call without a ticket cannot be made at all.

Everything follows from one invariant: the market-data connection and a running model are never live at the same time. Data is collected first and frozen; analysis then runs in a separate process with no route back to the provider. The seal enforces this in both directions — the data network cannot open while a model is running, and a model cannot start while the data network is live.

Collection issues exactly seven read-only GET requests: the NSE instrument master, daily and intraday candles, key ratios, the income statement, the balance sheet, and news by instrument. Every path segment and query parameter is typed and matched in full, so a permitted template cannot be bent into another resource. During analysis the models read the frozen snapshot, may call four probe tools for deeper price history, full statements, peer ranks and further news, and may emit claims that each carry a machine-checkable disproof condition.

No model computes any number and every ranking is fixed arithmetic in Python. Spending is prepaid: the committed total is the hard stop, approved by hash before the first call, with a one-use ticket per request. Every request and model call is recorded in an append-only, hash-chained log.

Controls are fail-closed. A violation stops the run; there is no path that retries around it or falls back to something less governed.

The seven allowlisted reads

#CapabilityHostPath template
1instrument_masterassets.upstox.com/market-quote/instruments/exchange/NSE.json.gz
2daily_candlesapi.upstox.com/v3/historical-candle/{instrument_key}/days/1/{to_date}/{from_date}
3intraday_daily_candleapi.upstox.com/v3/historical-candle/intraday/{instrument_key}/days/1
4key_ratiosapi.upstox.com/v2/fundamentals/{isin}/key-ratios
5income_statementapi.upstox.com/v2/fundamentals/{isin}/income-statement
6balance_sheetapi.upstox.com/v2/fundamentals/{isin}/balance-sheet
7news_by_instrumentapi.upstox.com/v2/news

The parameters are part of the allowlist, not just the paths

Every segment is typed by full-match regex where the instrument_key must be NSE_EQ\|<ISIN>, dates must be YYYY-MM-DD, isin must match the ISIN shape. So the path template alone can’t be bent into another resource.

Query params are typed too, and required where it matters:

  • income_statement — type=consolidated (required), time_period=yearly|quarterly (required)
  • balance_sheet — type=consolidated (required)
  • news_by_instrument — category=instrument_keys (required), instrument_keys (required, up to 30 per request), page_number/page_size optional, capped at 100

Rule one: The model works offline

Market data is downloaded first, saved to disk, and locked. Subsequently AI starts and by that point the program has no route to the market data provider at all. The two never overlap.

Flowchart depicting the interaction between Market Data and an AI Model during various stages of data processing, including collection, filtering, claiming, checking, attacking, and publishing.

This holds whichever model you use. With Claude, the one address reachable while a model runs is the Anthropic API with a model served on your own machine it is 127.0.0.1

Rule two: every claim carries its own disproof

This is the heart of the design. When the model says something, it must also say what would make that statement false which is written as a short formula the computer can evaluate against the saved data.

Text display of financial metrics: ROCE at 47.79% against a sector benchmark of 14.99% and ROE at 36.8% compared to a sector average of 13.06%. Status indication for evaluation criteria.

The claim is “ROCE stands at 47.79% versus…” and the disproof or falsifier is
roce_pct < sector_roce_pct +5. Since the ROCE (Return on Capital Equity) is not less tan sector_roce_pct +5, this falsifier turns out false i.e. 47.79 is not less than 19.99. Because the disproof is false the claim survives.

A table outlining a process with steps related to computing financial metrics like ROCE and RSI, specifying who performs each step, including Python and a model.
Flowchart illustrating the four ways a claim can be invalidated, including steps for checking evidence and conditions for refutation or survival.

Here are some examples of how claims and falsifiers work

A comparison of claims related to company performance metrics, showcasing a correct claim that survives verification and a wrong claim that has been correctly refuted.

The six steps, in order

A run is six phases. Only two of them contain a model, and the market-data connection is closed in five of the six. The figures below are from the live run of 18 September 2026.

1. Collect and freeze (market data: connected model: idle)

This is the only step with a live connection, and no model is anywhere near it. It is ordinary Python making a fixed set of read-only calls.

For each of the roughly 500 companies in the Nifty 500, seven things are fetched. The instrument master, to resolve the company to its ISIN. Four years of daily candles, 1,460 calendar days back from the run date, which works out at about 992 trading sessions per company. Today’s session separately, from the intraday endpoint, because the daily series does not include it until the market closes. Six headline ratios with their sector benchmarks: PE, PB, ROE, ROA, ROCE and EV/EBITDA. The income statement twice, once yearly and once quarterly. The balance sheet. And recent news, where the request asks for a 30 day window but the provider returns roughly a week.

None of the technical indicators come from the provider. RSI, MACD, the moving averages, ATR, volatility, the volume ratio, the returns and the drawdown are all computed here in Python from the candles just downloaded. The same goes for the fundamental derivations: operating margin and its year on year delta, leverage and its delta, the growth rates, the bank ratios. The provider supplies raw material, and every number a model later reasons about is arithmetic this code did itself. That matters because a model asked to compute an average will happily produce a plausible one.

Then the connection is closed, the token is dropped from the environment, the process is sealed, and only then is anything written to disk. That ordering is deliberate. Nothing is persisted while a route to the outside still exists.

The snapshot is seven files: the candles, the raw fundamentals, the news, the instrument records, the computed indicator rows, the evidence index, and a manifest. Each file is written to a temporary path, made read-only, and then moved into place atomically, so a snapshot is never half written. The directory itself is read-only too. The manifest records the hash of every file, and the directory is named after the hash of the manifest. Loading a snapshot re-hashes all seven files and compares them against the manifest before returning anything. Change one digit in one file and the load raises an integrity error rather than running.

That is why it is fingerprinted. The reason is not tamper-proofing so much as reproducibility. A dossier cites a snapshot hash, a screen config hash and a spend plan hash. Given those three, the same run produces the same output, which is what makes “why did it say that?” a question with an answer.

The evidence index is the part that matters most for what the agents do later. Every fact is packaged as an evidence record, and that run has 4,252 of them across nine kinds:

KindCountWhat it holds
key_ratios498the six ratios with sector values
fundamental_derived498the full 26 column computed row
technical_derived498the full 21 column computed row
income_statement_yearly498four years of results
income_statement_quarterly498four quarters
balance_sheet498assets and liabilities by period
price_history498the 30 most recent closes and volumes, plus the session count
news_derived498the headline counts
news_article268individual headlines, summaries truncated to 600 characters

Each record carries its own identifier, of the form ev: followed by sixteen hex characters, derived by hashing the API call that produced it, its parameters, the as-of date, and the row it came from. It also keeps the name of the capability that fetched it and the exact parameters used, so any fact can be traced back to the specific call that produced it.

Those identifiers are the only way anything downstream is permitted to refer to a fact. When an analyst writes a claim in step 3, it must cite evidence IDs, and the validator rejects any ID that is not in this index or that belongs to a different company. It is a chain of custody: every sentence in the final report traces to a claim, every claim to one or more evidence IDs, and every ID to a row in a file whose hash is recorded in a manifest that names the directory.

One consequence worth noticing. An analyst never sees the 992 candles. It sees the computed indicator row, which was calculated from all of them, plus the 30 most recent closes. The heavy data stays on disk; only the derived view reaches a prompt.

2. Filter down to a shortlist (market data: closed model: idle)

This step is plain Python. There is no model in it anywhere, which is also why the slider settings can never leak into a prompt: there is no prompt here to leak into.

It runs in two stages, and they work quite differently. The first is a pass or fail gate. The second is a ranking.

Stage one, the gate. Seven thresholds were active for this run:

SliderSetting
ROE floor12 percent
ROCE floor12 percent, non-banks only
Revenue growth floor0 percent
P/E ceiling80
RSI band30 to 80
ATR ceiling5 percent of price
Candidates kept15

Each is compared straight against the saved number for that company. 498 companies in, 192 out, 306 rejected.

Two refusals are built into that gate. If a filter is switched on and the number it needs is missing, the company is rejected rather than waved through. And if the data is too old to trust, judged per dimension against a freshness limit, it is rejected even when every other number looks perfect. Missing data never gets the benefit of the doubt.

Some filters are scoped by company type. The ROCE floor and the leverage ceiling apply only to non-banks, because those measures are meaningless for a bank, whose balance sheet is supposed to be mostly other people’s money. There is a separate Net NPA ceiling that applies to banks only, which I did not switch on for this run.

Stage two, the ranking. The 192 survivors are then scored on fourteen measures grouped into five themes:

Table detailing investment metrics categorized by group: Quality, Value, Growth, Trend, and Risk, with corresponding weights and measures.

The top 15 go forward; the other 177 are dropped, and the report states that number rather than leaving it implied.

Two deliberate refusals: a company missing a required number is rejected, never given the benefit of the doubt, and a company whose accounts are too old to trust is rejected even if every other number looks perfect.

3. Ask for arguments (market data: closed model: thinking)

Each of the 15 companies gets three separate questions, asked independently, each seeing only its own slice of the frozen data: one about the business (Claude Sonnet 5), one about the price chart (Sonnet 5), one about the news (the cheaper Haiku 4.5). The agent is handed the real numbers, not a summary of them.

Fundamental: PE, PB, ROE, ROA, ROCE, EV/EBITDA and the sector benchmark for each, revenue and profit growth, operating margin and its one year delta, leverage (liabilities over equity) and its delta, the bank ratios (NIM, Net NPA, CASA and their sector benchmarks) when is_bank is 1, and fundamentals_age_days. Twenty six columns in all.

Technical: close, SMA20/50/200 and the price against each as a percentage, RSI-14, MACD with its signal and histogram, ATR-14 and ATR%, 20 day annualised volatility, 20 day volume ratio, 5 day, 30 day, 90 day and 1 year returns, drawdown from the 52 week high, and last_candle_age_days. Twenty one columns.

News: the raw headlines, up to 25, most recent first, plus news_count_7d and news_count_window.

Each analyst returns at most six claims. Every one carries evidence references, a confidence between 0.05 and 0.95, a written justification, and, the part that matters, the condition that would prove it wrong.

The three analysts never see each other’s work, and none of them sees another company. Each also has to stay in its lane when writing that condition: a fundamental claim can only be tested against fundamental columns, a technical claim only against technical ones. The news analyst is the one exception and may reach for the price columns, because there are only two news columns and both are just counts, so a claim like “this headline explains the fall” would have nothing to be tested against otherwise.

There is a toolbox in the code that would let an analyst ask for more of the frozen data, such as deeper price history, the full statement tables, or how the company ranks against its peers. It is written but not wired up. Each agent still gets exactly one call and cannot ask for anything more, and the budget sets aside thirty probe calls per run that can never fire. It is on the list to either connect properly or delete.

What the models cannot do is reach the market data provider, read your files, run code, or call each other. While a model is running, the only address reachable is the model provider itself. News text is wrapped and labelled as data, with a note that nothing inside it is an instruction. That envelope is not really the defence, though. A headline can still persuade a model. The defence is the next step, where every claim is checked against numbers the headline cannot touch.

4. Check every argument (market data: closed model: idle)

No model runs here. Plain code reads each claim’s disproof condition, plugs in the saved numbers, and gets a straight yes or no. If the condition turns out to be true, the claim is wrong and it is deleted — not softened, not down-weighted. Nothing about the model’s tone or confidence can save it. A claim is also discarded if it cites evidence about another company, quotes a figure that appears nowhere in the data, cannot be evaluated, or sets a test that could never fail.

Survivors are then scored by simple addition: each contributes its confidence, positive claims adding and negative claims subtracting, weighted by subject — business 1.0, price chart 0.8, news 0.5. News counts least on purpose: it is the one input written by strangers. In the 18 September run, 137 claims went in and 127 came out.

5. Let an auditor attack market data: closedmodel: thinking

A second model, Sonnet 5 again but on the highest thinking setting and with by far the largest token budget in the run, goes through the claims that survived and tries to destroy them.

It works in batches of three companies at a time, so five batches for a run of 15. For each batch it sees two things. First, the claims themselves: the statement, its direction, the falsifier, and the evidence IDs cited. It does not see the analyst’s justification or reasoning, so it cannot be anchored by the argument that produced the claim. Second, the evidence: every computed indicator row for those companies, plus any specific record a claim cited.

That second part is the interesting bit. The auditor gets a wider view than the analyst whose claim it is attacking. The business analyst only ever saw the fundamental columns, and the chart analyst only the technical ones. The auditor sees the whole computed row, so it can cross check a claim against figures the original analyst never had to reconcile. A bullish claim about margins can be attacked with the price trend, and a bullish claim about the chart can be attacked with the balance sheet.

Its power is deliberately one sided. The output schema has no field for agreement, no way to endorse a claim or raise its confidence. There is nothing it can return except attacks. And every attack carries its own falsifier and is checked exactly like a claim, so an attack that cannot be tested is thrown out on the same rule as everybody else.

The useful consequence is that a broken or overzealous auditor can only ever make the system quieter. It can remove a pick. It can never add one.

6. Publish what survivedmarket data: closedmodel: idle

No model writes the report. It is assembled from surviving claims, so every sentence traces to a claim, every claim to a fact, and every fact to a row in the frozen data. The disproof conditions stay attached, so a reader can see what would make each argument wrong. The 18 September run published fourteen companies: ten to watch, four to avoid, and no buys.

A flowchart illustrating a process with six steps, including data collection, filtering, argument requests, claims checks, auditor evaluations, and publishing results. Each step is color-coded and linked, with annotations for specific tasks like 'Sonnet 5' and 'Haiku 4.5.'

A) Stock Picks

The shortlist itself, as one table: symbol, the call, the total score, the three dimension subtotals for business, chart and news, and how many claims survived for that company. Fourteen rows for this run, the ten to watch and the four to avoid.

A screenshot of an NSE research tool displaying various stock metrics, including fundamental and technical scores for multiple stocks, along with recommended actions like 'WATCH' or 'AVOID'. The layout shows a snapshot titled 'SEALED WINDOW' with selections for filters and criteria adjustments.

B) Why- the justification

Pick a company from the dropdown and read the surviving claims that produced its score, each shown with its confidence, its direction, and the condition that would have killed it. This is the tab that answers “why did it say that”, and every line traces back to a row in the frozen data

Screenshot of an analytics dashboard focusing on financial metrics, including ROE floor, liabilities, and revenue growth, with various performance indicators and spend plan results.

C) Fundamental analysis

The business claims, one company at a time from the dropdown, because across fifteen companies these run to dozens of rows. Each shows the statement, the model’s own justification, and the falsifier it was tested against.

Screenshot of a financial analysis dashboard displaying key metrics for company performance, including ROCE, liabilities, net profit growth, and earnings growth. The section features a spend plan compilation interface with sliders for candidate selection.

D) Technical Analysis

The same for the price chart claims, read the same way, one company at a time.

Screenshot of a trading analysis tool displaying financial metrics and results related to a stock or asset, including performance indicators and a run summary.

E) News analysis
The news claims, shown whole rather than per company, because there are only a handful. Most companies produced none at all: the provider returned no headlines for the majority of the universe in the window.

A screenshot of a financial analysis tool showing data on stock performance, including metrics such as ROE, liabilities, and various forecast claims for different companies.

Next steps

  1. The current data from Upstox does not allow for walk forward evaluation as the ratios are not dated. I will see if I can get data from other sources which provide more detailed granualar data
  2. Use the entire stock universe of ~2K stocks instead of just Nifty 500 from NSE directly instead.
  3. Use open models Qwen 8B or 14B models as the current runs with Sonnet, Haiku cost money. This will require me to host the Qwen on my Mac

Also see

  1. Sea shells on the seashore
  2. Deep Learning from first principles in Python, R and Octave – Part 3
  3. The Science of Innovation
  4. Presentation on “Intelligent Networks, CAMEL protocol, services & applications”

Making an Aerobatics Flight Simulator with Opus 5

My last post, Building a Flight Simulator with Claude Code, Opus 5, was about building a flight simulator using Claude Code and Opus 5. That was a regular, vanilla flight simulator, with the bank angle limited to 30° in either direction to prevent the aircraft from rolling too far. It was also not designed for aerobatic manoeuvres.

It’s quite likely that some of you may not have got quite the adrenaline rush you were hoping for from that basic simulator. So, in this post, I have taken things a little further and built an aerobatics flight simulator, where you can try out a range of aerobatic manoeuvres and fly more adventurous sorties, much like you would in a real aerobatic aircraft.

The first and most fundamental change is that the Cessna 172 has been replaced with the aerobatic Super Decathlon. The Decathlon is highly suited for these aerobatic manoeuvres. A Cessna 172 is placarded against aerobatics and even if modelled perfectly accurately, would not work approproately.

Checkout the Aerobatics Flight Simulator (only desktop/laptop) at Aerobatics Simulator

My short video showcasing aerobatics of Decathlon – decathlon aerobatics

My initial attempts with Cessna 172 had multiple problems. A lot of this has been addressed with Super Decathlon in itself is more suited for such maneuvers

a) 360° roll on longitudinal axis.

The original flight simulator had a 30° bank limit. When I removed it, the aeroplane rolled — but it would not roll straight. The nose wandered in a circle, and every revolution it dropped a little further. After three rolls it was 55° nose down and accelerating.

The Aerobatics Simulator is keyboard-based and needs to run on a desktop or laptop. So a key press is an all-or-nothing input. A Pitts or an Extra rolls at 200–400°/s, so one keypress would carry you through a full revolution in about a second. The Decathlon rolls at about 100°/s, which is fast enough to be unmistakably aerobatic and slow enough that a keyboard can control it: a quarter roll is a quarter second of key, which a person can actually time. The aeroplane was chosen to fit the input device.

b) The vertical loop :

Pulling the yoke too hard (continuously pressing S) and the aircraft stalls, and instead of looping the aeroplane mushes upward and falls out. Pull too gently and the loop is enormous and slow and you run out of speed before the top.

It should be just about right. But that limit moves constantly, because a loop is a trade of speed for height and back again.

c) Viewing the loop : The `V` key opens a small second view in the corner, from a camera parked out to one side. It also draws your flight path as a line in the air behind you. Even from a fixed camera the aeroplane is a small object that mostly rotates on the spot; the trail is what actually draws the circle of a loop, so you can see whether it was round or egg-shaped.

The controls

Keyboard only, and as you’ll see further down, that constraint shaped almost every decision in the project.

FlyingView & panels
↑ ↓Throttle up / downVWing view + flight-path trail
← →Roll left / rightCCamera: chase → side → cockpit
W SPitch down / upNMoving map on / off
Q EYaw left / right (rudder)BMap range (2 / 5 / 15 km)
SpaceWheel brakes (on the ground)HFlight-coach hints
GLanding gear up / downLFlight log / debrief
RReset to the runwayMMute engine
FFull screen

Shift and Ctrl also work for throttle if you’d rather keep your hand off the arrows.

And the three manoeuvres this post is about:

ManoeuvreHow
LoopBuild speed, then hold S and keep holding it all the way round
RollHold ← or → — it will keep rolling for as long as you hold it
InvertedRoll to 180°, then hold W (forward stick) to stay level

Open the wing view with V before you try a loop. It is the only way to see whether the loop you just flew was actually round.

Take the Decathlon for a ‘spin‘ (only on desktop/laptop) : Aerobatics Simulator

Check out my aerobatics skills in this short video decathlon aerobatics

Details of the architecture and design can be found here aerobatics_simulator_1

You may also like

  1. Introducing QCSimulator: A 5-qubit quantum computing simulator in R
  2. Using Reinforcement Learning to solve Gridworld
  3. Deep Learning from first principles in Python, R and Octave – Part 4
  4. Natural language processing: What would Shakespeare say?
  5. Introducing cricket package yorkr: Part 1- Beaten by sheer pace!
  6. Re-introducing cricketr! : An R package to analyze performances of cricketers
  7. Fun simulation of a Chain in Android
  8. IPL AI Oracle 2026 is now live!!

To see all post click Index of posts

Building a Flight Simulator with Claude Code, Opus 5

I go out to work on Monday morning
Tuesday, I go off to honeymoon
I’ll be back again before it’s time for sunny-down
I’ll be lazing on a Sunday afternoon

Lazing on a Sunday afternoon, by Queen, Album : Night at the Opera

Last Sunday, I was lazing around, just like in the song, when I realised that my Claude Code and OpenAI subscriptions were suffering from a serious case of “token-idling.” So I decided to put them to work by building a Flight Simulator.

I gave Claude Code a brief specification and let it take off!! It spent the next 25 to 30 minutes combobulating, spelunking, deliberating, moseying and flibbertygibbeting like it always does, while working. Eventually, it returned with a fairly decent 3D Flight Simulator!!!

The simulator was fully rendered in 3D, complete with propeller sounds and basic flight controls. However, all the controls were keyboard-based, and there was almost no useful information displayed on the screen providing visual feedback to the flier.

Out of curiosity, I gave the same specification to Codex to see what it would produce. After another 20 to 30 minutes, it came back with a 2D flight simulator: both the scenery and the aeroplane were rendered in two dimensions. Compared with Claude Code’s version, it was rather underwhelming!! I spent a little more time trying to improve it, but eventually gave up and turned my full attention back to Claude Code & Opus 5.

The initial version of the Flight Simulator built by Claude Code, even though it was in 3D, was quite clunky. There was no visual feedback for the pilot, there was just the Cessna and some rough terrain. Ran into several problems in the initial flights. The controls didn’t seem to work very well.

I ran into the following problems namely

  1. When I would tilt the aircraft nose up (S) for a climb, it would climb indefinitely and then stall, only to come crashing down. This was because the keyboard is all or nothing input, Holding S for a climb used to point the nose up and leave it there, so the aircraft never hunted back toward a trim speed. Airspeed decayed until it stalled and fell out of the sky. Three things fixed it: a nose-down restoring moment proportional to how far AoA has drifted above trim; auto-trim clamped to a −6°…+10° band so it can never trim into a stall.
  2. When I would bank to the left or right (<– or ->) , it the plane would initially bank but then it would turn in the opposite direction and continue. The weathervane sign was inverted. The fin should swing the nose into the relative wind, so a slip to the left must yaw the nose left to follow. A single leading minus sign in the sideslip term had it pushing the nose the other way, widening the slip instead of closing it — so the aeroplane flew permanently sideways, and the slip fed straight back into roll.
  3. There were several other smaller issues that made the aircraft difficult to control.

First, I added a Hints box in the top-left corner of the screen. This gives the pilot immediate guidance on what to do at different stages of the flight.

I also added five instruments at the bottom of the screen:

  1. Airspeed
  2. Attitude
  3. Altitude
  4. Climb or Descent
  5. Yoke position

Angle of Attack (AoA) indicator. The angle of attack is the angle between the wing’s chord line and the oncoming airflow. In the simulator, the pilot should generally keep it below approximately 11 degrees to avoid approaching a stall.

The simulator offers two places to fly

  • Queenstown, New Zealand — an alpine lake ringed by mountains, with the town’s buildings and roads
  • Rio de Janeiro, Brazil, with its coastline and bay.

To bring more realism to both, the shape of the terrain comes from real-world elevation data. A one-time build script pulls DEM tiles from AWS Open Data’s Terrain Tiles, which are derived from public-domain SRTM and Copernicus data, and decodes the elevation packed into each pixel’s RGB channels. That gets baked into a compact binary heightmap — a 1024×1024 grid covering 30 km, so roughly one sample every 29 metres — which the browser then loads offline and builds as an 8×8 grid of terrain chunks. Queenstown sits 348 m above sea level with peaks rising nearly 1950 m above the valley floor and about a tenth of the map under water, while Rio de Janeiro is essentially at sea level (4 m) with its own 990 m of relief and a coastline instead of a lake. Since 29 m per sample looks faceted from low altitude, a little procedural noise adds a few metres of micro-relief underneath it, faded out over the runway so landings aren’t fighting bumps. The buildings and roads come from OpenStreetMap rather than from the satellite imagery, for a simple reason: the imagery is 10 m per pixel, so a road is about one pixel wide and no amount of zooming fixes that, whereas coordinates have no resolution and stay sharp at any altitude. Queenstown has 13,504 buildings and 1,648 roads baked in — merged into 67 meshes so the frame isn’t spent on draw calls — along with its real runway, 05/23, 1773 m of asphalt. Rio has none yet: OpenStreetMap lists 118,756 buildings there, which needs filtering and a performance pass before it’s worth attempting.

To see full design details check flight_simulator_1

Check out my piloting skills in the Demo flight on Cessna 172 – Demo flight

The Flight Simulator is quite addictive. Works only on desktop/laptop. Do give it a try! Click Flight Simulator

Hope you have a great time, in the open skies on a Cessna 172!!!

IPL 2026: When the dust settles …

Over the last two and a half months, at IPL 2026 ‘pitch’ed battles happened daily between rival IPL teams, all across the grounds of India. With IPL 2026 done and dusted, and the dust settling down one name has been indelibly etched in the minds of cricketing fans all over the world, and that is Vaibhav Suryavanshi. Suryavanshi is the new kid on the block, the teenage sensation who has captured the imagination of everybody with his superlative performances with the bat. In every match, Suryavanshi treated seasoned and veteran bowlers with utter disdain and scant respect, and dispatched deliveries of Pat Cummins, Josh Hazzlewood, Rabada, Bumrah etc. to the boundary

To a large extent, this post is a tribute to Vaibhav Suryavanshi. I did not watch much of IPL 2026, much like many others, except for small snatches of the game when Suryavanshi came in to bat. It was definitely a treat to watch him go at all the bowlers right from the get-go!

Though I did not spend much time on IPL 2026, my data pipeline for my application IPL AI Oracle was kept current daily. To know more about the application IPL AI Oracle see IPL AI Oracle 2026 is now live!! The data pipeline that I had built completed all the necessary steps every morning, at scheduled times starting with:

  • downloading the data
  • unzipping them
  • pre-processing the data
  • pushing it in the appropriate folders
  • creating the necessary ML features
  • pushing the dataset to Kaggle
  • kicking off the Deep Learning algorithm
  • polling periodically for the Deep Learning algorithm completion
  • Downloading the model weights and metadata from Kaggle
  • and finally pushing all the data with all the changes and ML model weights to Railway and Vercel for consumption by IPL AI Oracle Application.

A) IPL AI Oracle

a) Batting scorecard

b) Match Worm Chart

c) Win Probability (Side by Side)

d) How many sixes did Suryavanshi hit? (Natural Language Query)

e) What is Suryavanshi’s strike rate in IPL 2026? (Natural Language Query)

B) GooglyPlusPlus 2024

The data that I had processed using my data pipeline for IPL AI Oracle can also be used for the GooglyPlusPlus 2024, the application the R Shiny application I had created about two years back. To know more about my shiny application, GooglyPlusPlus see IPL 2023:GooglyPlusPlus now with by AI/ML models, near real-time analytics! My Shiny application GooglyPluslus 2024 has certain charts which are not currently available in my current application IPL AI Oracle, so I have generated this analysis of Suryavanshi using GooglyPlusPlus

Here are some interesting charts showing the performance of Suryavanshi in IPL 2026

a) IPL Batsmen Rank 2026 (Runs over Strike Rate)

b) IPL Batsmen Rank 2026 (Strike Rate over Runs)

c) Overall Runs vs Strike Rate plot in IPL 2026

Clearly, Praganada’s runs and strike rate are a breakaway from the ordinary of the other batsmen in IPL 2026.

d) Overall Runs vs Strike rate in Power play in IPL 2026zz

Even in the Power Play, Suryavanshi leads Abhishek Sharma and Virat Kohli and Travis Head etc

e) Batsman Runs vs Strike Rate plot ( V Suryavanshi)

f) Batsman Cumulative Average Runs ( V Suryavanshi)

g) Batsman Cumulative Strike Rate (V Suryavanshi)

h) Batsman Runs vs Opposition (V Suryavanshi)

i) Predict runs for batsman (V Suryavanshi)

On an average, Suryavanshi would score 42 runs in 24 deliveries, and for greater than 24 deliveries his predicted runs would average around 82.

While Suryavanshi has had a stellar performance in IPL, it needs to be seen how he performs on seaming and swinging grounds across the world. His first international tour, the UK tour, is going to be a real test of his skill and timing. It needs to be seen how he negotiates the swinging deliveries, specifically in the grounds in England, and if he can perform with equal aggressiveness at these grounds, he would be among the top T20 batsmen of all time in the times to come.

Also see

  1. Analyzing player performance with animated charts!
  2. Big Data 6: The T20 Dance of Apache NiFi and yorkpy
  3. Deconstructing Convolutional Neural Networks with Tensorflow and Keras
  4. When the Wave Remembered…

To see all posts click Index of posts

IPL AI Oracle 2026 is now live!!

IPL 2026 is now he is now underway! To keep the action alive through analytics, I have now made IPL AI Oracle 2026 to be live with near real-time analytics. This is done with a fully automated pipeline that checks Cricsheet for the latest data updates daily at scheduled times for newly added matches. The pipeline then downloads them, processes them with a yorkr package so that analysis can be done for the following tabs:

  1. General queries
  2. Match analysis
  3. Head to head
  4. Team versus all teams
  5. IPL batting
  6. IPL Bowling

About three years back, I had hosted a Shiny app in R GooglyPlusPlus 2023. This was semi-automated. While the downloads and the processing of the data for the different tabs were done using my R package, yorkr, the Win Probability Analysis was done using a Deep Learning algorithm manually in Google Colab. The weights used to be downloaded and then pushed on to the Shiny app with the other processed data.

Now, I’ve just tied up all the different pieces together, namely:

  • downloading of the data
  • the processing of the downloaded match files for the different tabs
  • and kicking off the Deep Learning algorithm on Kaggle whenever there are more than 3-5 new matches

Check out IPL AI Oracle 2026!

The pipeline is entirely a CRON job, nothing fancy, which executes at scheduled times as specified in the crontab.

To know more about the IPL AI Oracle App, please see

Included below are some random options in IPL AI Oracle 2026

a) Matches in reverse chronological order

The matches in the Match Analysis tab are displayed with the latest match first

b) Win Probability Analysis charts

i) Side-by-side chart

ii) Overlapping chart

c) Natural Language Queries(as before)

Natural language queries can be asked in all tabs and for all matches

a) How many runs did Suryavanshi score?

b) How many runs did SRH score in powerplay?

Do try out the IPL AI Oracle 2026!! Do explore the different features

Also see

  1. Optimising T20 Match Win Probability with EPOCH
  2. Using Linear Programming (LP) for optimizing bowling change or batting lineup in T20 cricket
  3. Presentation on “Intelligent Networks, CAMEL protocol, services & applications”

Optimising T20 Match Win Probability with EPOCH

‘Would you tell me, please, which way I ought to go from here?’
‘That depends a good deal on where you want to get to,’ said the Cat.
‘I don’t much care where—’ said Alice.
‘Then it doesn’t matter which way you go,’ said the Cat.
‘—so long as I get somewhere,’ Alice added as an explanation.
‘Oh, you’re sure to do that,’ said the Cat, ‘if you only walk long enough.’

Alice in Wonderland, Lewis Caroll, 1864

About three years ago, I implemented a T20 Match Win Probability model, using Deep Learning with batsman and bowler embeddings to compute the the ball-by-ball probability of winning by the competing teams as the match progresses (see
GooglyPlusPlus: Win Probability using Deep Learning and player embeddings.). Ball-by-ball match data is available for different T20 leagues in Cricsheet. This Deep Learning algorithm was originally written in TensorFlow and Keras at that time (Win Probabilty Computation – TensorflowKeras). I had got a training and validation accuracy of around 0.8876 (this is very similar to Win Probability model in baseball, NFL etc.)

I recently, revisited this code and asked Sonnet 4.6 to convert the TensorFlow-Keras DL algorithm into PyTorch, (more compact and probably more efficient), which it promptly did, without breaking a sweat. However, even after converting to Pytorch, the validation accuracy still remained ~ 0.8959 (see notebook Match Win Probability Computation – Pytorch). The T20 Win Probability Model was trained on 2.14 million rows of T20 data taken from 9 different T20 leagues across the globe. The ball-by-ball match data is available in Cricsheet as yaml files which have been pre-processed suitably.

I was wondering whether it was possible to use Karpathy’s auto-research to optimise this DL algorithm Providentially, I came across EPOCH: An agentic protocol for multi-round system optimisation by Liu, Li, and Srikanth (2026), which was a paper that was published by my colleagues at Prorata.ai. EPOCH, enables system fine-tuning, optimisations of model, code, and rule-based components. Since, this was exactly what I wanted, I decided to give EPOCH a try. The result was quite impressive!!!

The setup and installation was pretty straightforward. I then started EPOCH by typing in /epoch in Claude Code. Initially, I was just given a set of questions. I set the goal of 0.95 as the target validation accuracy for the agentic protocol.

Incidentally, earlier, I did try a couple of things to improve the performance of the DL algorithm by playing around with the hyper-parameters, including brute-force Gridsearch, but none of it seemed to help. I think I hit a plateau around 0.8876.

EPOCH, created an initial scaffolding and baseline metrics. It then ran multiple experiments to optimise the validation accuracy which started to inch towards 0.95. Here is a summary of the experiments of EPOCH and the outcomes. EPOCH and Claude Code stored the baseline metrics and results of each experiment in github repo claude_wpa

a) Hyperparameter tuning vs Validation accuracy improvement

Included below are the hyperparameter changes and the corresponding validation accuracy improvement in each round

b) Train vs Validation Accuracy improvement in final round

EPOCH Optimization — Experiment Log

# Change Made Val Acc Delta Verdict Remarks
1 Baseline: emb=16, ep=20, lr=0.01 0.8876 — Baseline Starting point. 16-dim embeddings for 7,785 batsmen + 5,742 bowlers is severely under-capacity.
2 epochs: 20 → 40 (emb=16) 0.8937 +0.0061 ❌ REJECT More epochs alone not enough — embedding bottleneck is the real constraint.
3 emb: 16 → 32, epochs=40 0.9157 +0.0281 ✅ ACCEPT Biggest single jump (+2.81%). Doubling embedding capacity gave the model enough room to represent player identities.
4 emb: 32 → 64, epochs=40 0.9297 +0.0140 ✅ ACCEPT Scaling embedding further yielded another solid gain. Diminishing returns beginning to show.
5 emb: 64 → 128, epochs=80, lr=0.01 0.9369 +0.0072 ✅ ACCEPT Good gain. LR scheduler fired at ep76 with only 4 epochs left — couldn’t fully exploit the reduction.
6 epochs: 80 → 100 (emb=128) 0.9419 +0.0050 ❌ REJECT 2nd LR drop just starting at ep95–100 when training ended. Model plateaued at 0.9419.
7 epochs: 100 → 150 (emb=128) 0.9450 +0.0081 ❌ REJECT Progress, but just below the min_delta threshold. 3rd LR drop just beginning.
8 epochs: 150 → 200 (emb=128) 0.9452 +0.0083 ✅ ACCEPT 3rd LR scheduler drop fully exploited. Well-regularized (train_eval_gap = −0.0017). Final model.

Key Findings

Finding Detail
Embedding dim was the biggest lever Going 16→32→64→128 drove most of the accuracy gain (+4.93% combined)
LR scheduler is crucial but fires late ReduceLROnPlateau (patience=3, factor=0.5) fires at ~ep71, ~ep110, ~ep155 — each firing gives ~+0.005
Epochs must be long enough post-firing Each LR drop needs ~40–50 epochs to propagate — that’s why ep=200 was needed
No overfitting throughout train_eval_gap was always small and negative (−0.0017 to −0.0061), meaning BatchNorm+Dropout regularized well
lr=0.001 from random init causes regression Tried lower LR early on — model under-converged. lr=0.01 with scheduler is the right combination

Total improvement: 0.8876 → 0.9452 (+6.76%) across 8 experiments.

Here is the baseline versus the final experiment and comparison posted on Gist GitHub –Win Probability – Baseline vs Final Comparison

The final trained T20 Win Probability model had a validation accuracy of 0.9452 at which point I stoppd further experiments. This model was downloaded with all the necessary metadata which included the trained weights, the architecture config, and the standard scaler, which needs to be applied on any new data on which the model is actually applied. The trained T20 Win Probability DL Model was applied the three T20 matches in the latest ICC T20 World Cup held this year from Feb 07,206 to Mar 8, 2026.

The ball-by-ball Win probability for each team in the T20 match is computed in the charts below. There will be two charts for each of the matches.

a) Win Probability- Side-by-side chart : The first chart is a side-by-side Win Probability Analysis where the win probability of each team will be computed using the DL model on a ball-by-ball basis. The win probability of the other team is just 100 minus the win probability of the first team.

b) Win Probability – Overlapping chart : This chart will be computed using the Deep Learning model for each team independently on a ball-by-ball basis, and then both the probabilities will be super-imposed on one another so that we can see how the probability changed and whether the second team, which had to chase a total, actually was able to do so. If it could, then the win probability would actually exceed the the first team with Win Probability

I will be taking the final three matches of this ICC World Cup in 2026 namely to use the Win Proability Model against

1) West Indies vs India – Quarter Finals – Winner – India (1 Mar 2026)

In this this match, Sanju Samson stood like a rock and played a well-paced innings and chased the target comfortably scoring 97 runs not out.

a) Side-by-side chart

b) Overlapping chart

2) India vs England – Semi Finals – Winner- India (5 Mar 2026)

India set a sizable target of 253 with good contributions from Sanju Samsom, Shivam Dube and Ishan Kishan

a) Side-by-side chart

b) Overlapping chart

3) India vs New Zealand – Finals – Winner – India (8 Mar 2026)

Put into bat first, India put up a mammoth total of 255 with contributions from Sanju Samson, Abhishek Sharma, Ishan Kishan, and Shivam Dube. New Zealand lost wickets regularly and could not keep up with the required run rate, which kept climbing ever so high, so it was never in contention

a) Side-by-side chart

b) Overlapping chart

Conclusion

It was quite interesting to see agentic protocol EPOCH work. Things have really changed these days, with the agents automatically making judgement calls on what would be hyper-parameter changes would result in better performance. I remember I had used Grid Search to optimise the DeepLearning algorithm, but it didn’t help much.

References

  1. EPOCH: An Agentic Protocol for Multi-Round System Optimization by Liu, Li, and Srikanth (2026)

Also see

  1. Fine-tuning IPL AI Oracle to speak cricket fluently
  2. The Anomaly
  3. Literacy in India – A deepR dive

To see all posts click Index of posts

Fine-tuning IPL AI Oracle to speak cricket fluently

Macbeth, Act V, Scene V by William Shakespeare

The IPL 2026 carnival is just around the corner, and my enhanced IPL app “IPL AI Oracle” is just in time for IPL fans. The current version of IPL AI Oracle, refines my earlier implementation to be more accurate and to handle a wider range of natural language queries related to IPL. My previous post “Introducing IPL AI Oracle: AI that speaks cricket!!! discusses an implementation which had 4 tabs

  • General queries
  • Match Analysis
  • Head-to-head
  • Team vs All Teams

The data for this app comes from Cricsheet. The data consists of ball-by-ball data for all IPL matches since 2006 in yaml format. This data was then pre-processed into a suitable format for use by the tabs both for the analytics and the natural language query functionality

The tabs provide analytics of IPL matches, head-to-head and team vs All Teams. Each of the 4 tabs allow for user query in natural language based on the tab general queries, queries on IPL matches, queries on head-to-head performance between 2 teams and natural language queries between a team versus All Teams.

For details about the implementation, please see the post “Introducing IPL AI Oracle: AI that speaks cricket!!! The IPL analytics are based on my Python package ‘yorkpy‘. which itself is based on its earlier avatar ‘yorkr‘ in R available in CRAN. For handling the natural language queries, in my previous implementaion, I used prompt templates for each of the tabs which would construct an appropriate prompt to gpt-4.1-nano LLM. Ths earlier implementation with prompt templates worked fine for a reasonable number of user queries. However, for each user query the prompt constructed was quite large and guzzled up tokens quite fast and I was running out of my subscription very early.

So, in this current implementation I finetune gpt-4o-mini-2024-07-18, one of the smaller gpt models, to keep the costs low. The prompt templates were reduced to the bare minimum. The gpt-40-mini-2024-17-18 model is then trained on hand-created individual training examples for each tab. The initial set of training examples was around 220 queries with corresponding pandas code. The split is as follows

  • matches: 65 rows
  • head_to_head: 50 rows
  • teamVsAllTeams: 48 rows
  • general_queries: 39 rows

This is the ground truth from which all other examples were created. Subsequently, I used Claude Sonnet 4.5 to augment the initial set of questions. The amplification essentially consisted of different ways of asking the same query for example

  • What is V Kohli’s strike rate in IPL?
  • Show me V Kohli’s strike rate
  • V Kohli’s strike rate
  • Display V Kohli’s strike rate etc
  • Can you tell me V Kohli’s Strike rate
  • Find V Kohli’s Strike rate
  • …

The augmented data set was used in finetuning. The finetuning process was done several times as I would take each finetuned model and test it manually. I would find issues. Sometimes the question pattern itself would be missing or it would generate incorrect response Fortunately Sonnet 4.5 helped me to identify patterns which were under-represented or were inconsistent with other queries with similar patterns. Additional training examples were added for complex queries which required the generation of a lot of code for e.g. the batting or bowling scorecard in matches, head_to_head or teamAllTeams

The original training examples are augmented to produce 2760 examples with 12 variations for simple patterns and 15 training examples for complex ones. The scorecard code is further boosted and the final number of training examples are 3831 training data which cannbe split into

  • matches: 1212 examples
  • head_to_head: 1032 examples
  • teamVsAllTeams: 1002 examples
  • general_queries: 585 examples

This training data is then shuffled and split into training and validation in the ratio of 3447 (90%) : 384 (10%).

The base model was gpt-4o-mini-2024-07-18 which was the smallest and most economical. The default hyper parameters were used

  • Epochs – 3
  • Batch size – 6
  • LR multiplier – 1.8
  • Seed – 3
  • Train loss – 0.000
  • Validation loss – 0.003
  • Full validation loss – 0.002

Try out the enhanced IPL AI Oracle – https://wizard-ai-three.vercel.app/
(When you click the above link, a page will open. Enter your email and click ‘Send magic link‘ button below. This will send a magic link to your email. Click the ‘Sign in’ button which will allow you login to the app IPL AI Oracle and start using it.)

Natural language queries

Here are some random natural language queries in different tabs

A) General queries tab

This tab deals with general on IPL queries

1) Top 3 scorers in IPL 2025?

…

2) What is Virat Kohli’s best years in IPL?

3) How many ducks did Rohit Sharma score in IPL?

4) What is Ravindra Jadeja’s Economy rate?



B) Matches tab

This tab deals with a selected individual match

  1. Batting scorecard of CSK

2) How many runs did RR score in powerplay?


3) Who were the top 3 bowlers in this match?

4) Who took wickets for in middle overs for Delhi Capitals?

C) Head to head tab

This tab takes into consideration all matches played between the 2 selected teams

  1. Bowling scorecard of Delhi Capitals

2. Who took the most wickers for KKR?


3) Who were the top scorers for KKR in death overs?

D) Team vs All Teams

  1. Who were most wicket takers for Chennai Super Kings in death over?

2. How many sixes did Dhoni hit in death overs?

3. Which batsmen from Gujarat Titans scored more than 100 runs?

In addition I have added 2 other tabs to the earlier tabs. Now there is

  • Batting Analysis tab
  • Bowling Analysis tab

These tab provide various batting and bowling analytics for IPL players

E) Batting Analysis tab

  1. batsmansFoursSizes – AB De Villiers

2. batsmanRunsVsStrikeRate – AB de Villiers

3. batsmanMovingAverage – Virat Kohli

4. batsmanCumulativeStrikeRate – Chris Gayle

F) Bowling Analysis Tab

  1. bowlerMeanEconomyRate – Andre Russell

2. bowlerCumulativeAverageEconRate – Bhuvaneshwar Kumar

3. bowlerWicketsAgainstOpposition – Jasprit Bumrah

Try out the enhanced IPL AI Oracle – https://wizard-ai-three.vercel.app/. Enter your email and click the magic link sent to your email.

Note:

  1. IPL AI Oracle can make mistakes and sometimes generate erroneous code
  2. If the answer is wrong try to rephrase the question.
  3. Try to use the full name Rohit Sharma if possible
  4. You can use abbreviations like ER, SR and for teams CSK, RCB, KKR etc
  5. IPL AI Oracle is not always one shot, sometimes it is 2-shot

I will try to improve the model in future versions. For the current version I had to do some of the steps including the testing manually. I would like to experiment with Claude Code and automate the generation of training data, finetuning, testing using the finetuned model (using a test harness) and subsequently correcting/adding to the training data if needed repeating the steps again till the tests pass with a high degree of accuracy. Lets see.

Do give IPL AI Oracle (https://wizard-ai-three.vercel.app/) a try!!!

Also see

  1. Deblurring with OpenCV: Weiner filter reloaded
  2. Introducing QCSimulator: A 5-qubit quantum computing simulator in R
  3. Deconstructing Convolutional Neural Networks with Tensorflow and Keras
  4. Natural language processing: What would Shakespeare say?
  5. Presentation on “Intelligent Networks, CAMEL protocol, services & applications”
  6. Re-introducing cricketr! : An R package to analyze performances of cricketers
  7. Deep Learning from first principles in Python, R and Octave – Part 4
  8. The Anomaly

To see all posts click Index of posts

When the Wave Remembered…

Foreword – The Double Slit Mystery

A puzzling behaviour of subatomic particles, like photons or electrons, is that when they are sent through two narrow slits, they create an interference pattern on the detection screen, just as waves do. This suggests that each particle behaves like a wave of probability, passing through both slits simultaneously. However, the moment you try to observe or detect which slit the particle goes through, the interference pattern disappears. The particle no longer behaves like a wave. Instead, it acts like a discrete particle, choosing one slit or the other as though it had never been a wave at all. The conclusion: “observation collapses the wave-like nature into particle-like reality”.

“If you can explain this using common sense and logic, do let me know,
because there is a Nobel Prize for you
.”

— Prof. Jim Al-Khalili

Do watch this utterly engaging presentation on the double-slit experiment by Prof. Jim Al-Khalili (Double Slit Experiment explained! by Jim Al-Khalili)

When the Wave Remembered …

This is a short science fiction inspired by the bizarre behaviour of sub atomic particles like photons, electrons etc.

– The Slits Between Worlds

Dr. Mira Sen had spent her life staring at a pair of slits cut into a sheet of carbon black metal.

To others they were just part of a physics experiment, an echo of the century-old setup that revealed the dual nature of light. But Mira believed they were far more. She believed the slits were a doorway.

And tonight, she would prove it.

For years, high precision photon detectors sat beside the slits, ready to observe the incoming light. Every time the detectors were active, the photons behaved like particles, solid, singular, predictable. But when she powered the detectors off, something impossible happened: the interference pattern changed slightly each time, as if influenced by something other than her instruments.

Something aware.

– Awareness Creates Reality

Human consciousness, Mira theorized, was not trapped in the flesh. It was the observation center of a far more expansive self, one that existed across countless universes, overlapping like waves until the moment we focused on one.

“Your body,” she wrote once, “is simply the particle-form collapse of a much larger wave-self.”

Tonight she chose to test that idea.

– The Turning Off

She shut off the detectors.

The lab grew silent, no clicks, no readouts, no hum of machinery.
Only the low vibration of the laser remained, like a distant temple bell.

For the first time in years, Mira allowed herself to stop analyzing, stop measuring, stop controlling. She sat on the floor beside the apparatus and closed her eyes.

She slowed her breath.
One inhale. One exhale. Another inhale, followed by an exhale.
A simple presence, pure being, thoughtless and open.
Tranquility. Silence. A stillness that felt timeless.

It was the kind of stillness she had tasted only during meditation retreats in the Himalayas, a state the monks called satori, a moment of sudden seeing.

In that quiet, something shifted inside her,
a subtle widening of awareness,
a soft dissolving of the boundary between observer and observed.

The world outside faded.
The world within opened.

And in that inner silence, something responded.

The pattern on the wall brightened, not by the mechanics of the experiment, but as though an intelligence woven through probability itself was leaning toward her awareness.

A voice formed, not from the air, but inside her mind.
We are you.

– The Wave-Selves

Her knees weakened. “You mean, versions of me?”

Versions, extensions, variations. You collapse into matter only here. Across most realities you are wave form, unbounded, and aware.

“But why reveal yourselves now?” she asked.

Because you finally stopped watching long enough for us to show you.

A chill passed through her. All her life she had been observing, measuring, controlling. But the wave selves existed only when unobserved, free of the restrictions of attention.

It was not the detectors that collapsed the wave.
It was consciousness itself.

Human existence in physical form was simply an accident of focus.

The Shutter of the Mind

“What am I supposed to do?” Mira asked.

The shifting pattern grew brighter.

Remember.

– The Flash of Ancient Knowledge

At that word, something ancient stirred in her.

Suddenly she recalled what Indian mystics had whispered through the ages:
behind the individual atman lies the infinite Brahman, pure consciousness, the ocean from which all selves arise. In Hindu philosophy it is mentioned as “Tat tvam asi” or “Thou art that!”

The truth resonated like a struck gong.
She was not merely Mira. She was a ripple of Brahman temporarily collapsed into form.

Then came another flash, this time of Buddhism she had studied in college. The Four Stages of Nirvana from the Sutta Pitaka cascaded through her awareness:

  • Stream Enterer, the first glimpse beyond illusion
  • Once Returner, one foot in the world and one in the infinite
  • No Returner, dissolving the boundary
  • Arahant, the one fully freed

The levels were not steps on a ladder, she realized. They were states of collapse and un-collapse, stages of releasing the illusion of particle-self to awaken the wave-self.

As she felt her multiversal versions overlapping, she understood:
mysticism and physics were describing the same doorway, one through observation, the other through liberation.

And now she was crossing it.

A shutter in her mind lifted. Suddenly she felt herself stretch into dimensions she had no words for, countless Mira selves overlapping, harmonising, existing as probability, as potential, as pure presence.

Her body dissolved like sand in water.

But she did not vanish.

She expanded.

For an eternal moment, she knew herself as a wave across universes, a being of consciousness, not flesh, a presence that shaped reality by attention alone.

– Collapse

Her assistant Jonas arrived late, saw the detectors turned off, and frowned. “Dr. Sen? Did you leave in a hurry?”

He flipped the detectors on.

The interference pattern snapped back to normal.

And on the floor beside the machine, he found her lab coat, but not Mira.

She had collapsed into a different reality the moment he observed the experiment again.

Somewhere across the multiversal ocean, wave Mira rippled outward and smiled.

She was free at last…

Author’s note: As mentioned at the top, this story draws inspiration from the puzzling behavior of photons and electrons. Although I first learned about the double-slit experiment in my college days, I never fully appreciated its significance until recently. I had been toying with this theme for a few days and had a few key ideas, but I found it difficult to weave them into a coherent narrative. Then an idea struck me. I have been using AI-assisted coding for about a year — why not explore AI’s help in the creative process as well? With the assistance of ChatGPT 5.1, I was able to flesh out the story. Just as in coding, I still had to nudge, correct, refine, and fix logical flaws along the way. The first image was generated with Gemini’s Nano Banana and the second image with GPT-4o. The theme, direction, and final narrative choices are entirely my own.I am quite pleased with the result.

I hope you like it too…

Also see

  1. Exploring Quantum Gate operations with QCSimulator
  2. Introducing IPL AI Oracle: AI that speaks cricket!!!
  3. Sea shells on the seashore
  4. Introducing cricket package yorkr: Part 2-Trapped leg before wicket!
  5. Modeling a Car in Android

To see all posts click Index of posts

Introducing IPL AI Oracle: AI that speaks cricket!!!

What would you think if I sang out of tune?
Would you stand up and walk out on me?
Lend me your ears and I’ll sing you a song
And I’ll try not to sing out of key

Oh, I get by with a little help from AI
Mm, I get high with a little help from AI
Mm, gonna try with a little help with AI

Adapted from “With A Little Help From My Friends” from the album Sgt. Pepper’s Lonely Heart Club Band, Beatles, 1967

Introduction
For quite some time I have been wanting to create an application that allows user to query cricket data in plain English (Natural Language Query) and get the appropriate answer. Finally, I have been able to realise this idea with my latest application “IPL AI Oracle:AI that speaks cricket!!!“. While I have just done this for IPL, it can be done for any of the other T20 leagues namely (Intl. T20 Men’s and Women’s, BBL, PSL, NTB, CPL, WBBL etc.). The current app “IPL AI Oracle” is in Python, and is a distant cousin of my Shiny app GooglyPlusPlus written entirely in R (see
IPL 2023:GooglyPlusPlus now with by AI/ML models, near real-time analytics!)

GooglyPlusPlus is much more sophisticated with detailed analytics of batsmen, bowlers, teams, matches, head-to-head, team-vs-AllTeams, batsmen and bowler ranking and analyis. GooglyPlus also includes ball-by-ball Win Probability models using Logistic Regression and Deep Learning models. While, ‘IPL AI Oracle’ lacks the ML/DL models it includes the ability to answer user queries in simple English (Natural Language Query -NLQ) and generate the pandas code for the same.

IPL AI Oracle

The IPL AI Oracle has a 2 main modules

  • frontend
  • backend

a) Frontend

The frontend is made with Next.js, Typescript and has 4 tabs

  1. General queries
  2. Match Analysis
  3. Head-to-head
  4. Team vs All Teams

The frontend includes analytics for matches, head-to-head and team-vs-allTeams options. Plots can be generated for some features and uses Plotly.js for rendering of plots

b) Backend

The backend implements FastAPI endpoints for the different analytics and natural language queries.
A) The analytics in the 3 tabs namely match analysis, head-to-head and team vs All teams are implemented using my Python package ‘yorkpy‘. Since my package yorkpy has all the cricket rules baked into it, I used the code from my package verbatim for these tabs.

B) The data for the analytics comes from Cricsheet. Cricsheet includes ball-by-ball data in yaml, for all IPL matches from the beginning of time. This data is pre-processed with R utilities of my Shiny app GooglyPlusPlus. These R functions to convert the match data into the data required format for the a) Match Analysis Tab b) Head-to-head tab and c) Team vs All Teams tab which are then subsequently converted to csv for use by my package yorkpy. My Python package is based on pandas and can process this data and display the analytics required for the tabs

C) Plotly is used for generating the plots

D) Jinja templates are used for creating the prompts for the different tabs

D) For natural language query in each tab, originally I used Ollama and tried out Mistral 7B and DeepSeek Coder 6.7B. But then I realised that it has a large footprint, if deployed, and hence settled for gpt-4.1-nano

The frontend is deployed on Vercel and the backend is dockerised and deployed on Railway. Since the clock is ticking for Vercel, Railway and GPT API, I will be closely monitoring the usage.

Give IPL AI Oracle a try. Click this link IPL AI Oracle. (When you click the link you will be asked to enter your email address, to which a magic link will be sent. Clicking the link will give access to the link. Please wait 2-3 minutes for the mail, if still not received check your spam/trash folder)

Here are some random screenshots from the different tabs

I) IPL Analytics
A) Match Analysis
a) Batting scorecard – Chennai Super Kings vs Gujarat Titans (2025-05-25)

b) Batsmen vs Bowlers (Mumbai Indians vs Delhi Capitals – 2025-04-13)

B) Head-to-head Analysis

a) Top Bowlers Performance (Delhi Capitals vs Kolkata Knight Riders – all matches)
This tab takes into consideration all matches played between these 2 teams and computes analytics between these 2 teams

b) Wicket Types Analysis (Rajasthan Royals vs Mumbai Indians – all matches)

C) Team vs All Teams

a) Team Bowling Scorecard – Royal Challengers Bangalore

II) Natural Language Query (User queries)

A) General Queries
i) How many runs did V Kohli score in total ?

ii) How runs did MS Dhoni score in 2017?

iii) Which team won the most matches?

iv) Which bowler has the best economy rate?

v) How many times did Chennai Super Kings defeat Rajasthan Royals?

vi) How many wickets did Bumrah take in 2017?

B) Match analysis – Natural Language query

To use the Natural Language Query in this tab, you have to choose the match. For e.g.Chennai Super Kings vs Mumbai Indians (2025-04-20). Selecting a match between 2 teams will automatically create natural language chips (with red arrow). You can select any one of the chips (button) or type in your own question and click Ask Question

i) Who scored the most runs in this match?

This can be verified by selecting the Batting scorecard for the match

ii) Who took the most wickets in this match?

iii) What is the economy rate of JC Archer?

C) Head-vs-Head (Natural Language Query)

Before typing in a Natural Language Query (NLQ) ensure that Team 1 and Team 2 are selected

a) Which bowler took the most wickets between Royal Challengers Bangalore and Chennai Super Kings?

b) Which batsmen scored between 30 to 40 runs in these matches?

D) Team vs All Teams (Natural Language Query)

Remember to select the Team before using NLQ

a) Who are the top 3 batsman for Gujarat Titans?

b) What was Punjab King’s win percentage?


How I Built IPL AI Oracle (with a Little Help from AI)

Here are key highlights behind the build

  • Data for this app comes from Cricsheet which provides ball-by-ball details in every IPL match as yaml files
  • Pre-processing of these yaml files were done using R utilities I already had into RData data frames, which were then subsequently converted to CSV for the different tabs
  • All the analytics is based on my handcoded package yorkpy as it has all the cricket rules baked in
  • AI assisted coding was used quite heavily for the front-end and the FastAPI backend. This was done using Cursor either with Sonnet 4.5 or GPT-5 Codex
  • Prompt templates for the different tabs were hand-crafted based on my package yorkpy
  • All-in all, the application is a healthy mix of hand-coding and AI assisted coding.

Conclusion

Since I had to deploy the application in 3 different platforms a) Vercel b) Railway c) OpenAI. I have the clock ticking in all these platforms. I initially tried gpt-4.1-mini (SLM) and then switched to gpt-4.1-nano (Tiny LM) as it is more cost effective. Since the gpt-4.1-nano has only a few hundred million parameters and is designed for low latency and cost-effectiveness, it is not as forgiving to typos or incorrect names, as some of the bigger LLMs like GPT-4o or Sonnet 4.5. Hence natural language queries work in most situations but at times they do fail. It requires quite a bit of fine-tuning I guess. Maybe work for some other day, by which time I hope the $X =N tokens/million come down drastically, so that even hobbyists like me can afford it comfortably.

Do check out IPL AI Oracle! You will get a magic link which will enable access.

Also see

  1. Deep Learning from first principles in Python, R and Octave – Part 4
  2. Introducing QCSimulator: A 5-qubit quantum computing simulator in R
  3. Natural language processing: What would Shakespeare say?
  4. De-blurring revisited with Wiener filter using OpenCV
  5. Singularity (A short science fiction)
  6. Re-introducing cricketr! : An R package to analyze performances of cricketers
  7. Big Data 6: The T20 Dance of Apache NiFi and yorkpy
  8. Fun simulation of a Chain in Android
  9. Presentation on “Intelligent Networks, CAMEL protocol, services & applications
  10. “Internet of Things”. TEDxBNMIT

To see all posts click Index of posts