N
Noah Sheldon
BlogProjectsSkillsExperienceLinks
free ai audit
Home/Blog/Stakeholder Management/Future of AI Development: Trends and Predictions for 2025
OverviewStakeholder ManagementAI Trends2025

Future of AI Development: Trends and Predictions for 2025

Analyzing the shift towards small language models, local inference, and the rise of sovereign AI infrastructure.

Jul 19, 2026
12 min

TL;DR - The AI landscape in 2025 is defined by three shifts: models are getting smaller and cheaper, inference is moving to the edge, and agentic systems are replacing simple chatbots. For engineering leaders, the strategic question is not "which model do we use?" - it is "how do we build systems that work reliably, cost predictably, and evolve safely?" This post analyses the trends that matter for stakeholders making AI investment decisions, based on lessons from shipping AI products at a global ratings agency, xThreads, and PostQueue.


The Three Shifts Reshaping AI

If you follow AI news, every week brings a new model, a new benchmark, a new claim. For engineering leaders and stakeholders, the signal-to-noise ratio is terrible. Here are the three trends that actually change how you build products - filtered through the lens of what has worked (and what has not) in production.

The three shifts - model size, inference location, capability architecture

1. Model SizeBigger is betterSmaller + cheaperPhi-3, Llama-3, Mistral servedCost ↓ 90% vs. GPT-42. InferenceCloud-onlyEdge + on-deviceONNX, CoreML, WebGPULatency ↓ 20×, privacy ↑3. ArchitectureChatbotAgentic orchestratorPlan → Execute → Self-correctAutonomous task completion

Trend 1: Model Specialisation - Smaller, Cheaper, Better

The dominant narrative of 2023–2024 was "bigger models are better." GPT-4 set the standard and every competitor raced to match its parameter count. In 2025, that narrative has flipped. The most valuable models for production systems are not the largest - they are the ones that fit your cost budget, latency SLA, and domain.

Small language models (SLMs) like Phi-3, Llama-3-8B, andMistral-7B now handle 80% of production use cases at 10% of the cost of GPT-4. The remaining 20% - complex reasoning, creative writing, multi-step planning - still benefit from frontier models, but you route only those requests to the expensive endpoint.

// The routing pattern every production system should use
async function routeQuery(query: string): Promise<ModelEndpoint> {
  if (query.length < 100 && query.includes("summarise")) {
    return SLM_ENDPOINT    // $0.0001/token - good enough
  }
  if (query.length > 2000 || query.includes("code")) {
    return FRONTIER_ENDPOINT // $0.01/token - only when needed
  }
  return SLM_ENDPOINT
}
// This pattern saves 70–90% on inference costs in production

For stakeholders, the implication is clear: one model does not fit all. The smart investment is not in picking the best model - it is in building a routing layer that dispatches each request to the right model for the job. We use this pattern at a global ratings agency and it reduced our inference costs by approximately 80% while improving response times.

Cost comparison - SLM vs. frontier model for different task types

Summarisation$0.001Frontier: $0.01Complex reasoningSLM: failsFrontier: $0.01Classification$0.0005Frontier: $0.005Routing 80% of traffic to SLM = 80% cost reduction

Trend 2: Edge AI - Inference Without the Cloud

Sending every piece of data to a cloud API introduces latency, privacy, and costproblems. Edge AI - running models directly on the user's device - solves all three.

In 2025, every major browser supports WebGPU, which means you can run quantised models client-side using ONNX Runtime orllamafile. Apple's Core ML and MLX framework make local inference a first-class citizen on macOS and iOS.

For PostQueue.app, we moved content classification (spam detection, topic tagging, sentiment analysis) from a cloud API to an on-device model. The result:

  • Latency: 2 seconds (cloud) → 100 milliseconds (edge)
  • Cost: $0.003 per request → effectively $0 (one-time device model)
  • Privacy: Content never leaves the device
// Running a local model via ONNX runtime in the browser
import * as ort from "onnxruntime-web"

async function classifyLocally(text: string): Promise<Classification> {
  const session = await ort.InferenceSession.create("/models/classifier.onnx")
  const input = new ort.Tensor("float32", encode(text), [1, 128])
  const output = await session.run({ input })
  return parseClassification(output)
}
// No API call. No network. No cloud cost.

Key lesson - For stakeholders evaluating AI investments: every request that stays on-device reduces cloud dependency, improves user experience, and eliminates a data-privacy conversation. Edge AI is not just a technical optimisation - it is a trust and cost strategy.


Trend 3: Agentic Systems - From Answers to Actions

The most consequential shift in 2025 is the move from conversational AI toautonomous agentic systems. Instead of asking an AI to "write a summary," you give it a goal like "find all overdue accounts, send reminders, and update the CRM" - and it executes the entire workflow.

At a global ratings agency, we have deployed agentic systems that automate research workflows. An analyst gives a high-level brief, and the agent:

  1. Searches internal documents and external sources
  2. Cross-references financial data across filings
  3. Drafts a structured report with citations
  4. Flags contradictory data for human review
  5. Submits the draft for approval

This is not a chatbot answering questions. This is a digital workercompleting a multi-step process that used to take hours.

Agentic workflow - from analyst brief to deliverable

Analyst Brief1. Search & Collect2. Cross-Reference Data3. Draft Report4. Flag Contradictions5. Submit for ApprovalHuman Review Gate ✓Self-correcton error

For stakeholders, this is the most important trend to understand. Agentic systems are not like chatbots. They have autonomous execution, which means they need guardrails, observability, and escalation paths. The ROI is enormous - a single agent can replace hours of manual work per day - but only if you build the safety infrastructure to support it.


Trend 4: AI Safety and Governance

As AI systems move from "nice to have" to "mission-critical," the governance requirements scale proportionally. Financial services is ahead of most industries here - regulators already require model validation, audit trails, and explainability. But every industry is moving in this direction.

The practical implication for engineering leaders: build auditability into the architecture from day one. Every AI action should be logged. Every model decision should be traceable to a specific input. Every output should be capturable for post-hoc analysis.

// Every AI system should log this, by default
interface AIAuditRecord {
  timestamp: string
  modelVersion: string
  input: string           // truncated for PII compliance
  output: string
  latencyMs: number
  costCents: number
  guardrailHits: string[] // which safety rules were triggered
  humanReviewId?: string  // was a human involved?
  decisionPath: string[]  // which routing/triage decisions were made
}

// This record is not optional - it is your compliance defence

Key lesson - The teams that treat AI governance as a product feature - not a compliance checkbox - will be the ones that deploy AI fastest, because they will have the data to prove it works safely.


What This Means for Engineering Leaders

The trends above converge on a single strategic recommendation for stakeholders making AI investment decisions:

Route, Don't Replace · Edge First · Audit Always
  1. Build a routing layer. Do not commit to one model. Build infrastructure that can dispatch to different models - SLM, frontier, local - based on task requirements. This is the single highest-ROI engineering investment in AI today.
  2. Move inference to the edge. Start with content classification, content moderation, and simple generation. These tasks do not need cloud models and the latency+cost savings are immediate.
  3. Design for agency. The future of AI is not conversational - it is autonomous. Design your systems to expose tools and actions that agents can call. An API designed for chatbots will not work for agents.
  4. Govern from day one. Audit logging, model versioning, human-in-the-loop gates, and cost tracking are not afterthoughts. They are the infrastructure that makes deployment possible in regulated environments.
  5. Invest in observability. The hardest AI problems are not accuracy - they are latency spikes, cost explosions, and silent degradation. Monitor every dimension: response time, token usage, error rate, guardrail hit rate.

2025 is the year AI stops being a separate initiative and becomes embedded in how every product is built. The teams that treat it as infrastructure - not a feature - will be the ones that ship fastest, spend least, and sleep best at night.

Chapter

Stakeholder Management

Reading level

Overview

About the author

NS

Noah Sheldon

Applied AI/ML @ Fitch

NS

Noah Sheldon

Associate Director, Applied AI/ML @ Fitch Ratings

Get a free AI opportunity audit

never miss a deep dive

New essays land when they ship. No spam, unsubscribe anytime.