In my previous post Building an Agentic Stock Picker with Anthropic, I implemented an agentic stock picker built with Claude Opus 5. The application analyses NSE stocks under strict governance. The rule everything follows from is simple: the market-data connection and a running AI model are never live at the same time. Data is downloaded first and frozen on the laptop; only then do the agents run, with no route back to the data provider. In the agentic stock picker Sonnet writes the fundamental and technical claims besides acting as the auditor while Haiku handles the news.
Here the claim is a single checkable assertion about one company for e.g. a sentence citing figures from the frozen data (“MCX trades at a P/E of 55.00x against a sector P/E of 32.62x, so it is expensive on earnings”), paired with the tripwire that would prove it wrong. Since, the Claude runs cost money, I also wanted to explore Open Models.
As mentioned in my earlier post, I wanted to run the agentic stock picker on Open Models hosted entirely on my own machine, with no data sent anywhere and no cost per run. I tried Qwen, GLM, GPT-OSS, Mistral, Magistral, Gemma and Nemotron, and compared every one against Claude on the same set of stocks. No single open model matched Claude, so the post ends with a consensus of three open models from different labs: Google’s Gemma, Zhipu’s GLM and Alibaba’s Qwen.
Small models exposed weaknesses that a strong model had quietly papered over, and fixing them made the whole system better, Claude runs included.
Disclaimer:Note: This is not investment advice and has been done for purely research purposes.
The setup
- Machine: a MacBook Pro with an M3 Max chip and 64 GB of memory.
- Model server: Ollama, which runs open models locally and serves them on the laptop’s own network interface.
- Every model analysed the same 15 shortlisted stocks from the same frozen snapshot. The results were compared with Claude’s on that snapshot. Claude was used as the reference. I wanted to see if a local model can reach similar conclusions.
Two numbers do most of the comparing:
- Rank agreement asks whether the model puts the 15 stocks in the same order, best to worst, as Claude. 1.0 means an identical order; 0 means no relationship.
- Score gap asks whether the scores themselves are close: the average difference from Claude’s, as a fraction of how spread out Claude’s scores are. Lower is better.
How the agents work (a quick recap)
Each AI agent (fundamental, technical and news) looks at one company and writes claims, such as “MCX is expensive on earnings: P/E 55.00 against a sector P/E 32.62”. Every claim must carry its own falsifier/tripwire, which if evaluated to true would result in the claim being discarded. We can think of the falsifier/tripwire of the claim as the negation of the claim. For that claim, the tripwire/negation is “if P/E is below the sector’s P/E”.
Python then checks every tripwire against the frozen data:
- If the claim’s negation/tripwire is true, the claim is proven wrong and discarded.
- If the tripwire could never be true, the claim can’t be tested, so it’s discarded too.
- Only claims that could have been proven wrong, but weren’t, survive and count towards the score.
An auditor model then tries to knock down the survivors, and plain arithmetic turns what’s left into a score and a BUY, WATCH or AVOID.
Ten local models were tried on the same 15 stocks, and seven were set aside:
| Model | Lab | Rank agreement with Claude | Outcome |
|---|---|---|---|
| Gemma 4 31B | 0.87 | Member, the strongest local model | |
| Gemma 4 26B | 0.74 | Good and faster, but from the same family as the 31B | |
| GLM-4.7-Flash | Zhipu | 0.71 | Member, fast (about 25 minutes) |
| Qwen 3.8 27B | Alibaba | 0.57 | Member; weaker alone, but the average is improved by it |
| Magistral | Mistral | 0.33 | Dropped; the average was made worse by adding it |
| Mistral Small 3.2 | Mistral | 0.28 | Dropped; TITAN was called a BUY |
| Nemotron 3.5 Lightning | NVIDIA | 0.05 | Dropped; 12 of 15 stocks were rated BUY |
| gpt-oss 20B, Qwen 3 8B / 14B | OpenAI, Alibaba | (earlier data) | Dropped; too optimistic, or superseded |
Claude’s run (claude-sonnet-5 with claude-haiku-4-5, about $1.92 for 15 stocks) was used as the reference throughout.
1. Tripwires pinned to the stock’s own value
Qwen 8B kept setting a tripwire’s threshold at the stock’s current value, copied from the data. TITAN’s 30-day return was −4.112031%, and the model wrote:
“The price has underperformed … by over 4 percent.” Falsifier: if
return_30d_pct > -4.112031
Since the return is exactly −4.112031, the tripwire (falsifier) can never go off, so the claim can never be proven wrong. The checker correctly discards it as untestable. But 29 of the 35 pinned tripwires were on claims against a stock, so the negative claims were hit far harder than the positive ones. In the end, only 24 of 72 negative claims survived, against 70 of 92 positive ones, and every stock’s score was pushed upward.
2. Bigger model
Qwen 14B works better and does not pin negation on exact values
Qwen 14B was overly optimistic and one-sided. It wrote 95 claims in favour of stocks and 48 against, and found none of the four stocks Claude marked AVOID. On the stock MCX it praised strong returns and growth, all true, but never mentioned a P/E of 53 against a sector average of 32. This was corrected by also checking on the downside risks. After this the claims swung from 95 for and 48 against to 38 for and 64 against, with 7 AVOIDs where Claude had 4.
3. Claims written backwards
Claim: “P/E is below the sector P/E” (cheap on earnings) Tripwire: wrong if
pe <= sector_pe
Here the claim and the tripwire are in the same direction which is incorrect. The tripwire is actually the correct test for the opposite claim, “expensive”. Python evaluated it faithfully: is 70.7 ≤ 6.0? Since the tripwire did not fire a false “cheap” claim survived as a point in the stock’s favour. Three such claims helped push ACUTAAS to a BUY.
4. The signal table
The fix for claims written backwards was to stop asking the model to invent rules. A python based signal table is created that defines 26 indicators with 56 rules. Each indicator has a rule for owning the stock and a rule against, and each rule comes with its exact tripwire/negation.
| Indicator | FOR | AGAINST |
|---|---|---|
| P/E vs sector | below the sector: “cheap on earnings” | above the sector: “expensive on earnings” |
| ROCE vs sector | above the sector | below the sector |
| Debt trend | liabilities-to-equity falling | rising |
| RSI-14 | ≤ 30 oversold; 30–45 in an uptrend (a pullback); 55–70 bullish | ≥ 70 overbought; 30–45 in a downtrend |
| MACD | MACD line above its signal line | signal line above the MACD line |
The signal table supplied a fixed, round threshold for every rule, so a stock’s exact reading can no longer become a tripwire. Automated tests prove that every tripwire is the exact opposite of its rule, including at the edges, where mistakes like > versus >= or AND versus OR hide. The table also carries 105 worked examples from real stocks, always written figures first, conclusion after: “GMRAIRPORT trades at a P/E of 128.31x against a sector P/E of 56.75x, so it is expensive on earnings.”
Sometimes the Local LLM auditors turned out to delete true claims. Some auditors refuted by copying the claim’s own, negation minus a clause. GLM-4.7-Flash’s auditor argued:
5. Auditor changes
“ANANDRATHI’s P/E (78.50x) is higher than the sector P/E (49.79x), contradicting the claim of being expensive.”
That argument restates and confirms the claim rather than contradicting it. I added two rules:
- A challenge that merely restates the claim can’t remove it. This is checked logically, by testing whether the challenge’s condition follows from the claim’s.
- A comparison Python has already computed (78.50 > 49.79) can only be challenged with other evidence.
These changes were replayed on stored runs and these rules now blocked 7 of GLM’s 9 wrong vetoes, and none of Claude’s.
6. Weighting
The local LLM models still rated the shortlist well above Claude. The reasons
Valuation counted no more than a price trend.
Every model gave almost every claim 0.85–0.95 confidence, whatever the evidence. A bigger model didn’t help: Qwen 27B was exactly as confident as Qwen 14B.
So confidence for rule-based claims is now computed from the data alone for e.g. 0.55 when a indicator is just past the rule’s threshold, up to 0.90 when it’s well beyond. A P/E of 25 against a sector P/E 24 gets 0.58, similarly a P/E of 53 against sector P/E 32 gets 0.90.
Now the weight is assigned based on the group to which the indicator belongs
| Group | Weight | Counts for the stock | Counts against the stock |
|---|---|---|---|
| Valuation | 1.5 | P/E, EV/EBITDA or P/B below the sector’s | P/E, EV/EBITDA or P/B above the sector’s |
| Debt (leverage) | 1.25 | Liabilities-to-equity fell over the year (not banks) | Liabilities-to-equity rose over the year |
| Returns on capital | 1.0 | ROCE, ROE or ROA above the sector’s (ROCE not for banks) | ROCE, ROE or ROA below the sector’s |
| Growth | 1.0 | Revenue or net profit up on the previous year | Revenue or net profit down on the previous year |
| Margins | 1.0 | Operating margin up over the year | Operating margin down over the year |
| Bank quality (banks only) | 1.0 | Net interest margin above the sector, net NPA below it, or CASA above it | The reverse of any of these |
| Momentum | 1.0 | RSI-14 ≤ 30 (oversold); 30–45 while above the 200-day average (a pullback in an uptrend); 55–70 (bullish). MACD above its signal line | RSI-14 ≥ 70 (overbought); 30–45 while below the 200-day average. MACD below its signal line |
| Risk | 1.0 | Within 10% of the 52-week high; volatility below 20%; ATR below 2% of price | More than 20% below the 52-week high; volatility above 35%; ATR above 4% |
| Trend | 0.75 | Above its 200-, 50- or 20-day average; MACD above zero | Below those averages; MACD below zero |
| Price returns | 0.5 | Up over 30 days, 90 days or a year | Down over those periods |
| Volume | 0.5 | A 5-day rise on rising volume | A rise on thinning volume, or a fall on rising volume |
The Stock Picker approach
The whole pipeline, as it runs today, whichever model is plugged in.

Stage 1: Filters (pass or fail)
Each of the roughly 500 Nifty 500 stocks must pass every filter:
| Filter | Setting |
|---|---|
| Return on equity (ROE) | at least 12% |
| Return on capital employed (ROCE) | at least 12% (not applied to banks) |
| Revenue growth | not shrinking |
| P/E | at most 80 |
| RSI-14 | between 30 and 80 |
| Daily price range (ATR) | at most 5% of the price |
| Data freshness | results under 400 days old, prices under 5 days |
Missing data counts as a fail. Around 190 stocks usually pass. Debt isn’t a filter here; it’s judged later, by its trend.
Stage 2: Ranking to a shortlist of 15
The survivors are ranked on 14 measures in five groups, each measure converted to a percentile against the whole market rather than tested against a fixed cut-off:
| Group | Weight | Measures |
|---|---|---|
| Quality | 30% | ROCE vs sector, ROE vs sector, margin trend |
| Value | 20% | P/E, P/B and EV/EBITDA vs sector (lower is better) |
| Growth | 20% | revenue growth, profit growth |
| Trend | 20% | price vs 200-day average, MACD, 90-day return |
| Risk | 10% | ATR, volatility, distance from 52-week high |
A measure a company can’t have, such as ROCE for a bank, scores at the median rather than the bottom. The top 15 go forward.
Stage 3: The AI analysts
For each of the 15, Python works out which of the 56 rules hold.
Three analysts per stock. Each of the 15 shortlisted stocks is examined by three separate analysts. Each analyst is a separate call to the AI model, and each sees only its own kind of evidence:
| Analyst | Looks at | Rules from the signal table |
|---|---|---|
| Fundamental | valuation, returns on capital, growth, margins, debt, bank measures | 13 indicators, 26 rules |
| Technical | price trend, momentum, returns, risk, volume | 13 indicators, 30 rules |
| News | recent headlines about the company | none; its claims are written freely |
Each analyst chooses up to six claims, including both sides whenever both exist, and writes the argument. The news analyst reads the headlines separately. Python checks every claim, and the auditor challenges the survivors under the rules above.
The 56 rules. The rules are kept in a Python file called the signal table. For each of 26 indicators, one rule is given for owning the stock (FOR) and one against it (AGAINST).
A few rows of the 26 rules used by the fundamental analyst a few are shown below
| Indicator | FOR | AGAINST |
|---|---|---|
| P/E vs sector | cheap on earnings | expensive on earnings |
| EV/EBITDA vs sector | cheap on cash profit | expensive on cash profit |
| P/B vs sector | cheap on book value | expensive on book value |
| ROCE vs sector (not banks) | returns on capital above the sector | below the sector |
| ROE vs sector | returns to shareholders above the sector | below the sector |
A few of the 30 rules used by the technical analyst below
| Indicator | FOR | AGAINST |
|---|---|---|
| RSI-14 | oversold (≤ 30); pullback in an uptrend (30–45, above the 200-day average); bullish (55–70) | overbought (≥ 70); weak in a downtrend (30–45, below the 200-day average) |
| Price vs 200-day average | long-term uptrend | long-term downtrend |
| Price vs 50-day average | medium-term uptrend | medium-term downtrend |
| Price vs 20-day average | short-term strength | short-term weakness |
| MACD vs its signal line | MACD bullish | MACD bearish |
Stage 4: Score and recommendation
Each surviving claim contributes:
± dimension weight × group weight × confidence
Fundamental claims weigh 1.0, technical 0.8 and news 0.5; the group weights are in the table above.
For example, a stock clearly expensive on P/E (53 against 32, confidence 0.90) loses 1.0 × 1.5 × 0.90 = −1.35. Clearly strong ROCE earns 1.0 × 1.0 × 0.90 = +0.90. A clear 200-day uptrend earns 0.8 × 0.75 × 0.90 = +0.54. One clear valuation concern outweighs a strong return on capital.
The recommendation rules:
| Label | Rule |
|---|---|
| BUY | score ≥ 1.0, and positive claims in at least two dimensions, and no surviving claim against the business at confidence 0.6 or more |
| AVOID | score ≤ −0.5 |
| WATCH | everything else |
A single clear negative about the business blocks a BUY, and most high-quality companies are expensive, so BUY is rare by design.
The results for the seven open models against Claude
| Model | Maker | Rank vs Claude | Score gap | BUY / WATCH / AVOID | Run time |
|---|---|---|---|---|---|
| Claude (API, about $1.90 per run) | Anthropic | — | — | 1 / 9 / 5 | 14 min |
| Gemma 4 31B | 0.87 | 0.45 | 1 / 13 / 1 | ~1.5 h | |
| Gemma 4 26B | 0.74 | 0.46 | 1 / 12 / 2 | ~35 min | |
| GLM-4.7-Flash | Zhipu | 0.71 | 0.69 | 3 / 11 / 1 | 27 min |
| Qwen 3.8 27B | Alibaba | 0.57 | 0.69 | 2 / 11 / 2 | 1 h 35 min |
| Magistral | Mistral | 0.33 | 0.87 | 2 / 12 / 1 | 1 h 3 min |
| Mistral Small 3.2 | Mistral | 0.28 | 0.76 | 2 / 13 / 0 | 1 h 18 min |
| Nemotron 3.5 Lightning | NVIDIA | 0.05 | 1.50 | 12 / 3 / 0 | 20 min |
gpt-oss 20B (OpenAI’s open-weight model) was tested on the earlier data only. It tended to be too optimistic, with rank agreement 0.31, and wasn’t carried it forward. Neither did the Qwen 3 8B and 14B models, which the table above already covers.
- Google’s Gemma 4 31B agreed with Claude best by a clear margin, and wrote no invalid claims at all. It had been released only days before these runs.
- Some models are simply optimistic. Nemotron rated 12 of the 15 a BUY, including TITAN, the stock Claude rated worst. gpt-oss and Mistral leaned the same way to a lesser degree. What set them apart was a tendency to choose supporting facts over warning signs, even though every fact they were given was true.
- Speed and quality are not the same axis. The fastest models (Nemotron, gpt-oss) were among the least aligned, while GLM-4.7-Flash and Gemma 4 26B were both fast and reasonably close.
- Every model agreed on one stock. GESHIP (Great Eastern Shipping) was a BUY for Claude and every open model tested on this data. It had a P/E of 5.7 against a sector 15, returns well above its peers, falling debt and an upward trend.
- Several models agreed on the other end. TITAN was an AVOID for Claude, GLM, Qwen 27B and Gemma 4 26B. Its business is excellent (ROE 32%, revenue up 45%), but at 27.6 times book value against a sector 3.3, with rising debt and fading momentum, the valuation and debt weights pull it down.
A consensus of models
Because no single open model matches Claude, I tried averaging several. Models from different labs tend to make different mistakes:
| Combination | Rank vs Claude | Score gap |
|---|---|---|
| GLM + Qwen 27B | 0.80 | 0.57 |
| GLM + Qwen 27B + Gemma 4 31B | 0.82 | 0.52 |
The GLM and Qwen pair beat both of its members, and adding Gemma 4 31B improved it further. The trio still falls short of Gemma 4 31B on its own (0.87), which argues for giving Gemma a strong voice. So a three-lab consensus (Zhipu, Alibaba and Google) is taken to decide which stock gets a BUY or AVOID only when most models agree. The two Gemma models together scored higher still (0.92), but coming from the same family, they’re more likely to share blind spots.

Here is what we get when we run each LLM independently
| Stock | Claude | GLM-4.7-Flash | Qwen 3.8 27B | Gemma 4 26B | Gemma 4 31B |
|---|---|---|---|---|---|
| IKS | WATCH (+4.95) | WATCH (+3.34) | WATCH (+4.65) | WATCH (+2.85) | WATCH (+4.65) |
| GESHIP | BUY (+4.48) | BUY (+4.66) | BUY (+4.19) | BUY (+5.00) | BUY (+5.79) |
| ENGINERSIN | WATCH (+2.63) | WATCH (+3.25) | WATCH (+1.06) | WATCH (+0.36) | WATCH (+3.00) |
| LODHA | WATCH (+1.84) | BUY (+2.44) | WATCH (+2.21) | WATCH (+1.03) | WATCH (+1.85) |
| HINDZINC | WATCH (+1.83) | WATCH (+1.59) | WATCH (+3.75) | WATCH (+1.67) | WATCH (+2.50) |
| GABRIEL | WATCH (+1.56) | WATCH (+2.96) | WATCH (+0.28) | WATCH (+0.66) | WATCH (+1.56) |
| HBLENGINE | WATCH (+0.74) | WATCH (+2.49) | WATCH (0.00) | WATCH (+0.96) | WATCH (+0.63) |
| MCX | WATCH (+0.36) | WATCH (+0.88) | WATCH (+1.18) | WATCH (−0.17) | WATCH (+0.88) |
| LLOYDSME | WATCH (+0.30) | WATCH (+0.49) | WATCH (+1.76) | WATCH (−0.13) | WATCH (+1.57) |
| LUPIN | WATCH (−0.49) | WATCH (+1.51) | BUY (+2.90) | WATCH (+0.07) | WATCH (+1.27) |
| PREMIERENE | AVOID (−0.55) | WATCH (+1.94) | WATCH (+0.98) | WATCH (+1.87) | WATCH (+0.73) |
| MUTHOOTFIN | AVOID (−0.91) | BUY (+3.25) | AVOID (−1.51) | WATCH (+0.23) | WATCH (−0.13) |
| KALYANKJIL | AVOID (−1.07) | WATCH (+1.57) | WATCH (+2.28) | AVOID (−1.15) | WATCH (+1.09) |
| GLENMARK | AVOID (−1.39) | WATCH (−0.01) | WATCH (+0.30) | WATCH (−0.10) | AVOID (−0.89) |
| TITAN | AVOID (−2.28) | AVOID (−0.92) | AVOID (−0.75) | AVOID (−1.85) | WATCH (+0.40) |
| BUY / WATCH / AVOID | 1 / 9 / 5 | 3 / 11 / 1 | 2 / 11 / 2 | 1 / 12 / 2 | 1 / 13 / 1 |
| Rank agreement with Claude | — | 0.71 | 0.57 | 0.74 | 0.87 |
| Score gap vs Claude | — | 0.69 | 0.69 | 0.46 | 0.45 |
The number in brackets is each stock’s score. A stock is BUY at a score of 1.0 or more, but only if it is positive in at least two of the fundamental, technical and news areas and has no strong negative business claim. It is AVOID at −0.5 or below, and WATCH otherwise.
The Consensus strategy
The whole pipeline is run separately by each AI model, and a score and a BUY, WATCH or AVOID are produced for every stock . The finished results are then combined by regular Python.
- The order of the stocks is set by a weighted average of the models’ scores.
- The label is decided by a weighted vote on the models’ own labels. A stock is made a BUY only if BUY was given by models holding at least half the weight, and an AVOID if AVOID was given by models holding at least half.
- A BUY is blocked by a heavyweight objection. If AVOID is given by any model holding 0.3 or more of the weight, the stock cannot be made a consensus BUY.
- Everything else is labelled WATCH.
| Model | Weight of its vote, without Claude | Weight of its vote, with Claude |
|---|---|---|
| Claude | not in the vote | 0.40 |
| Gemma 4 31B | 0.4 | 0.25 |
| GLM-4.7-Flash | 0.3 | 0.175 |
| Qwen 3.8 27B | 0.3 | 0.175 |
| Total | 1.0 | 1.00 |
- Order: the weighted average of the scores
weighted score = Σ (model weight × model’s score)
With Claude, for GESHIP: 0.40 × 4.48 + 0.25 × 5.79 + 0.175 × 4.66 + 0.175 × 4.19 = +4.79.
2. . The BUY test
A stock is made a consensus BUY only when all three checks are passed:
- No veto: AVOID was not voted by any single model whose own weight is 0.3 or more (rule 3).
- Majority: the BUY weight is 0.5 or more.
- Score: the weighted score is +1.0 or more.
3. The veto: a BUY is blocked by a heavyweight objection
If AVOID was voted by any one model whose weight is 0.3 or more, the stock cannot be made a BUY, however large the BUY weight.
Who can cast a veto is therefore set by the weights:
| Version | Weight of 0.3 or more | Can veto a BUY on its own |
|---|---|---|
| Without Claude | Gemma (0.4), GLM (0.3), Qwen (0.3) | all three |
| With Claude | Claude (0.40) only | Claude only; Gemma (0.25), GLM and Qwen (0.175) can vote against a BUY, but cannot block one alone |
4. The AVOID test
A stock that is not a BUY is made an AVOID when either check is passed:
- Majority: the AVOID weight is 0.5 or more, or
- Score: the weighted score is −0.5 or below.
Unlike BUY, AVOID can also be reached through the score alone
5. Everything else is WATCH
This includes an exact even split, where the BUY weight and the AVOID weight are both 0.5. When the models are divided, “watch’ should be fine.
Using weights (without Claude)**
| Gemma (0.4) | GLM (0.3) | Qwen (0.3) | BUY weight | AVOID weight | Veto? | Result | Example |
|---|---|---|---|---|---|---|---|
| BUY | BUY | BUY | 1.0 | 0 | no | BUY | GESHIP |
| BUY | BUY | WATCH | 0.4 + 0.3 = 0.7 | 0 | no | BUY | illustration |
| WATCH | BUY | WATCH | 0.3 | 0 | no | WATCH: 0.3 is short of 0.5 | LODHA |
| BUY | BUY | AVOID | 0.7 | 0.3 | yes: Qwen’s 0.3 | WATCH | illustration |
| WATCH | AVOID | AVOID | 0 | 0.3 + 0.3 = 0.6 | — | AVOID | TITAN |
| WATCH | BUY | AVOID | 0.3 | 0.3 | — | WATCH: no majority either way | MUTHOOTFIN |
| AVOID | WATCH | WATCH | 0 | 0.4 | — | WATCH: 0.4 is short of 0.5, score −0.27 | GLENMARK |
| WATCH | WATCH | WATCH | 0 | 0 | — | WATCH, despite a score of +4.26 | IKS |
The results on 28 Sep data (with prices to 25 Sep 2026)
A) Without Claude (Gemma 0.4, GLM 0.3, Qwen 0.3)**
| Stock | Gemma 4 31B | GLM-4.7 | Qwen 27B | Weighted score | Consensus | Claude (not voting) |
|---|---|---|---|---|---|---|
| GESHIP | BUY (+5.79) | BUY (+4.66) | BUY (+4.19) | +4.97 | BUY | BUY |
| IKS | WATCH (+4.65) | WATCH (+3.34) | WATCH (+4.65) | +4.26 | WATCH | WATCH |
| HINDZINC | WATCH (+2.50) | WATCH (+1.59) | WATCH (+3.75) | +2.60 | WATCH | WATCH |
| ENGINERSIN | WATCH (+3.00) | WATCH (+3.25) | WATCH (+1.06) | +2.49 | WATCH | WATCH |
| LODHA | WATCH (+1.85) | BUY (+2.44) | WATCH (+2.21) | +2.14 | WATCH | WATCH |
| LUPIN | WATCH (+1.27) | WATCH (+1.51) | BUY (+2.90) | +1.83 | WATCH | WATCH |
| GABRIEL | WATCH (+1.56) | WATCH (+2.96) | WATCH (+0.28) | +1.60 | WATCH | WATCH |
| KALYANKJIL | WATCH (+1.09) | WATCH (+1.57) | WATCH (+2.28) | +1.59 | WATCH | AVOID |
| LLOYDSME | WATCH (+1.57) | WATCH (+0.49) | WATCH (+1.76) | +1.30 | WATCH | WATCH |
| PREMIERENE | WATCH (+0.73) | WATCH (+1.94) | WATCH (+0.98) | +1.17 | WATCH | AVOID |
| HBLENGINE | WATCH (+0.63) | WATCH (+2.49) | WATCH (0.00) | +1.00 | WATCH | WATCH |
| MCX | WATCH (+0.88) | WATCH (+0.88) | WATCH (+1.18) | +0.97 | WATCH | WATCH |
| MUTHOOTFIN | WATCH (−0.13) | BUY (+3.25) | AVOID (−1.51) | +0.47 | WATCH | AVOID |
| GLENMARK | AVOID (−0.89) | WATCH (−0.01) | WATCH (+0.30) | −0.27 | WATCH | AVOID |
| TITAN | WATCH (+0.40) | AVOID (−0.92) | AVOID (−0.75) | −0.35 | AVOID | AVOID |
B) With Claude (Claude 0.40, Gemma 0.25, GLM 0.175, Qwen 0.175)**
| Stock | Claude | Gemma 4 31B | GLM-4.7 | Qwen 27B | Weighted score | Consensus |
|---|---|---|---|---|---|---|
| GESHIP | BUY (+4.48) | BUY (+5.79) | BUY (+4.66) | BUY (+4.19) | +4.79 | BUY |
| IKS | WATCH (+4.95) | WATCH (+4.65) | WATCH (+3.34) | WATCH (+4.65) | +4.54 | WATCH |
| ENGINERSIN | WATCH (+2.63) | WATCH (+3.00) | WATCH (+3.25) | WATCH (+1.06) | +2.56 | WATCH |
| HINDZINC | WATCH (+1.83) | WATCH (+2.50) | WATCH (+1.59) | WATCH (+3.75) | +2.29 | WATCH |
| LODHA | WATCH (+1.84) | WATCH (+1.85) | BUY (+2.44) | WATCH (+2.21) | +2.02 | WATCH |
| GABRIEL | WATCH (+1.56) | WATCH (+1.56) | WATCH (+2.96) | WATCH (+0.28) | +1.58 | WATCH |
| LLOYDSME | WATCH (+0.30) | WATCH (+1.57) | WATCH (+0.49) | WATCH (+1.76) | +0.91 | WATCH |
| LUPIN | WATCH (−0.49) | WATCH (+1.27) | WATCH (+1.51) | BUY (+2.90) | +0.90 | WATCH |
| HBLENGINE | WATCH (+0.74) | WATCH (+0.63) | WATCH (+2.49) | WATCH (0.00) | +0.89 | WATCH |
| MCX | WATCH (+0.36) | WATCH (+0.88) | WATCH (+0.88) | WATCH (+1.18) | +0.73 | WATCH |
| KALYANKJIL | AVOID (−1.07) | WATCH (+1.09) | WATCH (+1.57) | WATCH (+2.28) | +0.52 | WATCH |
| PREMIERENE | AVOID (−0.55) | WATCH (+0.73) | WATCH (+1.94) | WATCH (+0.98) | +0.47 | WATCH |
| MUTHOOTFIN | AVOID (−0.91) | WATCH (−0.13) | BUY (+3.25) | AVOID (−1.51) | −0.09 | AVOID |
| GLENMARK | AVOID (−1.39) | AVOID (−0.89) | WATCH (−0.01) | WATCH (+0.30) | −0.73 | AVOID |
| TITAN | AVOID (−2.28) | WATCH (+0.40) | AVOID (−0.92) | AVOID (−0.75) | −1.11 | AVOID |
**Please note
Disclaimer:The stocks listed here should not be used as investment advice. This has been done for research purposes.
Some findings
- The clear calls are agreed by both. GESHIP is rated BUY and TITAN AVOID in both versions.
- More caution is shown by the version without Claude. Where the models disagree, WATCH is assigned.
- The weak end is mostly sharpened by adding Claude.
- Only 15 stocks on one date have been tested. The data is at least 2 weeks old
- This has been done for research purposes alone and is not a stock recommendation
See also
- Building a Flight Simulator with Claude Code, Opus 5
- Introducing IPL AI Oracle: AI that speaks cricket!!!
- Introducing QCSimulator: A 5-qubit quantum computing simulator in R
- Experiments with deblurring using OpenCV
- Singularity
To see all posts click Index of Posts





























































































