Amazon Web Services bundles dozens of AI and machine learning products under one roof. For developers, the hard part is not finding a service—it is choosing the right one without overpaying or building a Rube Goldberg pipeline.
This guide maps the AWS AI services landscape as of 2026: what each major offering does, when to reach for it, and how the pieces connect in typical architectures.
The three layers of AWS AI
Think of AWS AI in three layers, from highest abstraction to deepest control:
- Ready-made AI APIs — send text, images, or audio; get labels, translations, or transcriptions back
- Foundation model platforms — access and fine-tune LLMs and multimodal models through unified APIs
- ML infrastructure — train, deploy, and operate custom models on managed compute
Most teams start at layer 1 or 2 and only descend to layer 3 when off-the-shelf models cannot meet accuracy, latency, or compliance requirements.
Amazon Bedrock: generative AI hub
Bedrock is AWS's primary surface for generative AI. It provides access to foundation models from Anthropic, Meta, Mistral, Amazon, and others through a single API. You choose a model, send prompts (or multimodal inputs where supported), and pay per token.
Bedrock fits when you need:
- Chat, summarization, or code generation without hosting GPUs yourself
- Knowledge bases for RAG over S3 documents with managed chunking and retrieval
- Agents that call AWS Lambda functions or APIs based on model reasoning
- Guardrails for content filtering and PII handling at the API boundary
Bedrock is usually the first stop for LLM features on AWS. Fine-tuning support varies by model; check current model cards before assuming you can train custom weights.
Tip: Separate your prompt templates from application code. Version them like any other config so you can roll back behavior without redeploying services.
Amazon SageMaker: full ML lifecycle
SageMaker is AWS's end-to-end machine learning platform. It is broader than generative AI—it covers traditional tabular models, computer vision, and custom NLP training.
Core SageMaker capabilities:
- Studio — notebook-based development environment
- Training jobs — managed distributed training on CPU/GPU instances
- Model registry and endpoints — deploy models with auto-scaling inference
- Feature Store — share ML features across teams
- Pipelines — CI/CD for ML workflows
Use SageMaker when you need custom model training, proprietary architectures, or strict control over inference hardware. Bedrock handles generative tasks faster; SageMaker handles "we own the model weights and training data" requirements.
SageMaker JumpStart offers pretrained models and notebooks for common tasks—a middle ground between Bedrock's API and training from scratch.
Pre-trained language APIs
These services solve specific NLP tasks without you managing models:
Amazon Comprehend
Comprehend analyzes text for sentiment, entities, key phrases, language detection, and PII. It supports custom classification and entity recognition with your labeled data.
Reach for Comprehend when you need high-volume, low-latency text analysis and do not need generative output. Ticket routing, moderation queues, and document triage are classic use cases.
Amazon Translate
Machine translation for dozens of languages. Pairs well with Comprehend for multilingual support workflows.
Amazon Transcribe and Transcribe Medical
Speech-to-text for general audio and clinical settings (with appropriate compliance review). Often upstream of Comprehend or Bedrock for call center analytics.
Amazon Polly
Text-to-speech. Used in IVR systems, accessibility features, and content narration.
Vision and document AI
Amazon Rekognition
Image and video analysis: object detection, facial analysis, content moderation, celebrity recognition, and text in images (OCR). Common in media workflows and physical security integrations.
Amazon Textract
Extracts text, tables, and form fields from scanned documents and PDFs. Finance and logistics teams use Textract to automate invoice processing before feeding structured data into downstream systems.
For complex document Q&A, teams often combine Textract extraction with Bedrock knowledge bases or custom RAG.
Speech and conversational AI
Amazon Lex
Builds conversational interfaces (chatbots and voice bots) with intent recognition and slot filling. Integrates with Lambda for fulfillment logic. Lex V2 improved NLU capabilities, though many teams now embed Bedrock directly into custom chat UIs.
Amazon Connect + Contact Lens
Contact center platform with analytics powered by Transcribe and sentiment analysis. Relevant if your AI requirements sit inside customer support operations rather than product features.
Specialized and emerging services
Amazon Q — AWS's AI assistant for developers (code suggestions, AWS documentation Q&A) and business users (enterprise search over connected data). Distinct from embedding Bedrock into your product, though both may use similar models under the hood.
HealthLake and HealthScribe — healthcare-specific data stores and clinical documentation generation. Only relevant with appropriate regulatory oversight.
Forecast, Personalize, Fraud Detector — ML for time series, recommendations, and fraud. Not generative AI, but frequently grouped under "AWS AI" in architecture discussions.
Neuron and Inferentia/Trainium — custom silicon for cost-optimized inference and training. Infrastructure choices when SageMaker endpoints need better price/performance at scale.
Choosing a service: decision guide
| Need | Start here |
|---|---|
| Chatbot or copilot feature | Bedrock (+ optional knowledge base) |
| Sentiment on 10M reviews/month | Comprehend |
| Custom fraud model on tabular data | SageMaker Autopilot or custom training |
| Extract tables from PDF invoices | Textract → Lambda → database |
| Moderate uploaded images | Rekognition |
| Fine-tune an open-weight LLM on private data | SageMaker or Bedrock fine-tuning (model-dependent) |
| Real-time translation in app | Translate |
When two services overlap—Comprehend sentiment vs. Bedrock prompt for sentiment—compare cost per request, latency, and whether you need explainable structured output vs. free-form analysis.
Architecture patterns that work
RAG over private documents: S3 → Bedrock Knowledge Base (or OpenSearch + custom chunking) → Bedrock model → API Gateway/Lambda → client app.
Call center summarization: Connect → Transcribe → Comprehend or Bedrock summarization → CRM via Lambda.
Computer vision pipeline: S3 upload trigger → Rekognition → DynamoDB metadata → alert via SNS.
Custom model serving: SageMaker training → model registry → endpoint with auto-scaling → Application Load Balancer.
Keep IAM permissions tight. AI services often need S3 read access, KMS decrypt, and VPC endpoints if you avoid public internet egress.
Cost and operations
AWS AI pricing models differ:
- Per API call or per token (Comprehend, Bedrock, Textract)
- Per inference hour (SageMaker endpoints)
- Per training hour (SageMaker training jobs)
Set billing alarms early. Bedrock token usage can spike with verbose prompts or unbounded agent loops. SageMaker endpoints incur charges even at zero traffic unless you use serverless inference or tear down endpoints after hours.
Log prompts and outputs carefully—retention policies should respect privacy regulations. CloudWatch metrics plus model invocation logging help debug quality regressions when you change prompts or models.
Security and compliance
- Use private VPC endpoints for Bedrock and SageMaker when data cannot traverse the public internet
- Enable encryption at rest for S3 buckets holding training data and KMS keys you control
- Apply IAM least privilege—inference roles should not have blanket S3
*access - Review Bedrock guardrails and output filters before exposing generative features to end users
- Document model provenance for regulated industries—know which foundation model version processed which customer data
Getting started this week
- Enable Bedrock model access in the AWS console for your region (model availability varies)
- Run a Comprehend sentiment job on a sample CSV in S3 to baseline a classification task
- Deploy a Lambda + Bedrock hello-world that answers questions over a single PDF uploaded to a knowledge base
- Estimate monthly cost at expected request volume before committing in production
FAQ
Is Bedrock replacing SageMaker?
No. Bedrock optimizes access to foundation models. SageMaker remains the platform for custom training, proprietary models, and granular MLOps.
Can I use OpenAI models on AWS?
Not natively through Bedrock. Bedrock hosts AWS-supported foundation models. You can call external APIs from Lambda, but that is a separate integration pattern.
Which region should I use?
Pick the region closest to users for latency, but confirm your required models and services are available there—Bedrock model lists are region-specific.
Do I need ML expertise?
API-layer services (Comprehend, Textract, Bedrock with simple prompts) require software engineering skills more than ML research skills. Custom SageMaker training still benefits from ML experience.
AWS AI services reward clarity about the problem first. Pick the smallest service that satisfies the requirement, measure quality on real data, and only add complexity when metrics—not hype—demand it.
Further Reading
Discover more articles on similar topics across our network
Ventilator Vanguard: AI-Powered MultiOrganFailure Survival Engine Using AWS
Cubed




Comments
Loading comments…