AI Deploy Network
Investment BankingAI AgentAI Case Study

Agentic stock intelligence platform with a custom MCP orchestration layer

AI Deploy Network Approved

DigitalxCode needed defensible daily equity coverage across two dozen buckets without scaling headcount. Built a GPT-4o agent over a custom MCP orchestration layer, shipped as five containerised services on GCP Compute Engine, running unattended every night behind a live dashboard.

NavinBy Navin
Chapter 01

The Business Challenge

The business situation the organisation faced before the project began.

Equity research coverage does not scale with headcount. Producing a defensible daily view across two dozen equity buckets means repeating the same retrieval, cross-checking and write-up work every night, and an LLM asked to do it in one pass produces fluent output with no traceable basis. Two further problems made a naive pipeline unusable: the model had no memory of what it concluded yesterday, so it restated the same thesis as if it were new, and nothing in the loop could distinguish a grounded signal from a confidently hallucinated one before it reached the dashboard.

Chapter 02

Why AI Was the Right Solution

Why AI was the right approach — and what alternatives were considered.

The task is genuinely language-shaped: read heterogeneous market commentary, weigh it against a bucket thesis, and produce a written signal with reasoning attached. A deterministic screener can rank on numbers but cannot read a narrative and say why a thesis changed.

Alternatives considered. A pure quantitative screener was rejected because it cannot explain a thesis or incorporate qualitative catalysts, which is most of what a reader wants from a daily note. A single-pass LLM prompt per bucket was built first and then rejected: output quality was unverifiable and the system had no memory across runs. A fixed scripted pipeline with an LLM only at the write-up step was rejected because which buckets need deep retrieval genuinely varies with market conditions, and a fixed sequence handled that badly.

The decision was to keep the LLM for reasoning and language, move every consequential action behind a discrete inspectable tool, and put validation in the write path so the model could not persist an unchecked claim.

Chapter 03

How the Solution Was Delivered

Discovery, design, development, testing and rollout — the journey, not the tooling.

  1. 1Framed the nightly job as a tool-using agent rather than a prompt chain
  2. 2Built a custom MCP orchestration layer exposing four discrete tools
  3. 3Implemented bucket generation and RAG retrieval as separate, independently testable tools
  4. 4Added an LLM health validation tool and placed it before any database write
  5. 5Built a ChromaDB memory pipeline on text-embedding-3-small for cross-session context
  6. 6Containerised the system as five services and deployed to GCP Compute Engine
  7. 7Scheduled nightly generation and comparison jobs with APScheduler
  8. 8Shipped a live dashboard over the persisted signals
Chapter 04

Key Technical & Architecture Decisions

Architecture, model selection, workflow and trade-offs.

The central decision was to expose capability as four discrete MCP tools (bucket generation, RAG retrieval, LLM health validation, database write) and let the agent sequence them rather than hard-coding the order. Market conditions vary, and which buckets need deep retrieval is not knowable in advance.

The second was to put validation in the write path rather than on a monitoring dashboard. A validator that runs after persistence tells you that you shipped a bad signal. A validator between the model and the database stops it.

The third was memory as retrieval, not context stuffing. Prior conclusions are embedded with text-embedding-3-small into ChromaDB, so the agent retrieves what is relevant to tonight's bucket instead of carrying an ever-growing transcript.

Deployment used Docker Compose across five services (FastAPI, Next.js, PostgreSQL, Redis and ChromaDB) on a single GCP Compute Engine instance rather than managed services. The nightly workload is predictable, and one instance keeps both cost and operational surface small.

Chapter 05

Challenges & How They Were Solved

Obstacles hit along the way and how they were overcome.

The hardest problem was repetition. Early runs restated the previous night's thesis in new words, which erodes a reader's trust faster than being wrong does. Solved by making prior conclusions retrievable and having the agent reason about change against them rather than generating from a blank slate.

The second was ungrounded confidence. A model asked for an investment signal will always produce one, with equal fluency whether or not retrieval returned anything useful. Solved by making health validation a tool the run must pass through before the write tool, so an unsupported signal fails at the boundary instead of being surfaced.

The third was orchestration cost and latency. Letting an agent freely sequence tools invites loops. Bounded by scoping each tool narrowly and keeping the nightly job batched per bucket, so the blast radius of a bad sequence is one bucket rather than the whole run.

Chapter 06

User Adoption & Change Management

How users responded — training, change management and feedback loops.

This is an internal production system rather than a broad organisational rollout. Output is consumed through a live dashboard, with nightly comparison jobs so a reader sees what changed since the previous run instead of re-reading a full report.

No adoption-volume or user-productivity claim is made. Usage was not instrumented, and that omission is the single reason this record reports delivered system behaviour rather than a confirmed business outcome.

Business Impact

Business Outcomes & Impact

The measurable outcomes from the underlying AI Deployment — with the story behind them.

Hours Automated
0
Cost Savings
Currency not specified
Revenue Impact
Currency not specified
Measurable Business Outcome

Daily investment-signal coverage across 24 equity buckets produced autonomously, replacing per-bucket manual research cycles with a scheduled unattended run, with LLM health validation enforced before any signal is persisted.

Business Outcome Categories

Productivity ImprovementRisk ReductionAI Performance Improvement

The story behind the numbers

Daily investment-signal coverage across 24 equity buckets is produced autonomously with no human in the loop, replacing per-bucket manual research cycles with a single scheduled unattended run. Every persisted signal passes LLM health validation before it reaches the database, and cross-session memory means each run reasons about change rather than restating prior conclusions.

The headline number is 24 buckets a night at zero human touches. The number that actually mattered during development was how many signals never reached the dashboard. Moving validation from a post-hoc check into the write path changed the system's failure mode: instead of surfacing a confident but ungrounded signal and relying on a reader to catch it, the pipeline now fails closed. Coverage and trustworthiness were traded against each other deliberately, and the validation gate was chosen over raw throughput.

This is a builder-reported technical and operational outcome. No client-confirmed hours saved, cost saving or revenue impact is claimed, because usage was never instrumented.

Reflection

Lessons Learned

What surprised the team, what worked well, and what would be done differently.

  • An agent that cannot remember yesterday will confidently repeat it; memory is a correctness feature, not a convenience
  • Validation belongs between the model and the datastore, not on a dashboard beside it
  • Discrete, narrowly scoped tools make an agent's behaviour inspectable in a way one long prompt never is
  • Retrieval-based memory scales better than growing context windows and stays cheaper per run
  • A predictable nightly workload does not need managed autoscaling; five containers on one instance was the right call
  • Letting the agent sequence tools is worth it only where the work genuinely varies
  • Instrument adoption from day one; not having usage metrics is the only reason this case study cannot state a confirmed business outcome
Looking Ahead

Future Opportunities

Where this solution could go next.

The clearest next step is outcome instrumentation: recording analyst time per bucket before and after, so the operational gain can be stated as a confirmed figure rather than described.

Beyond that, backtesting the signal stream against realised returns would give the validation gate a quality score rather than a pass or fail. Coverage could widen past 24 buckets once per-bucket cost is measured. And a human escalation path for buckets where validation fails repeatedly would surface a persistent retrieval gap to a person, instead of silently producing no signal.

For Peers

Professional Reflections

Guidance for another builder tackling a similar problem.

If you are building a nightly agent, decide early what the system does when it has nothing good to say. Most designs implicitly answer produce something anyway, which is the worst option and the one you get for free. Put the gate in the write path and let a run come back empty.

Treat memory as retrieval over prior conclusions rather than a longer prompt.

And be honest about where the agency is earning its keep. Four tools with a model choosing the order was justified here because retrieval depth varies by bucket. A good share of agent architectures would be better served by a script, and that is worth measuring before committing to it.