AI/ML Solution Architect
Pratham InternationalAI/ML Solution Architect
Education & Assessment Platforms
About Pratham International
Pratham International is a US 501(c)(3) established to support learning innovations for children and youth in contexts across the globe. We enable the adaptation and spread of proven and scalable education solutions pioneered in India by Pratham Education Foundation to contexts across the global south through programming and technical assistance tailored to local realities. Today, Pratham innovations reach millions of children annually through partnerships across ~30 countries.
Through the Pratham Shah PraDigi Innovation Centre (PraDigi), Pratham International also seeks to create technology innovations for low-resource communities, and support models aimed at learning for school, work, and life for children and youth. We work on several cutting-edge projects in the domain of Generative AI, Automatic Speech Recognition, and Natural Language Processing.
The team’s innovative technology products include the PadhAI: Reading Assessment App, which uses advanced speech recognition to evaluate children’s reading fluency in 8 Indian languages and 6 global languages; Anytime Testing Machine (ATM), an AI-enabled question generation and grading system that allows learners to be assessed on any topic, anytime; and a suite of AI-powered chatbots designed for diverse learning contexts: from early childhood care, to youth employment readiness, and OCR-based analysis tools that digitize and assess student work at scale.
At Pratham International, we build for the long term. Whether it’s ensuring products are truly accessible across rural India and the global south through languages, building datasets to ensure context and accuracy for products, or gathering feedback from the field. Backed by partners like Google.org, Anthropic, and the Gates Foundation, we are committed to building digital public goods that ensure every child is in school and learning well.
Role Overview
About the work
We design and deliver AI systems that governments and large education programmes can actually run: curriculum-aligned item generation, paper assembly, handwriting OCR, rubric grading, and teacher-in-the-loop review. The work is already live as practice assessment in India and is being built as a national exam platform in Rwanda, with further country programmes behind it.
We are hiring a Solution Architect who sits between discovery and delivery. You will turn messy workshop notes and ministry requirements into architecture a team can build, secure, cost, and operate — RAG and multi-agent systems, LLM-as-judge pipelines, and the cloud/LLMOps layer underneath.
You will
- Own solution design from first workshop to a delivery-ready blueprint (services, data stores, trust boundaries, cost, ops).
- Architect LLM pipelines: generation, retrieval over curriculum knowledge graphs, standards/duplication/quality judges, constraint-based paper assembly.
- Design human + AI workflows with audit, segregation of duties, cycle-pinned config, and no student PII in model calls.
- Choose and govern models, prompts, eval gates, and fallbacks (including OCR: Azure primary, secondary fallback).
- Shape the Azure/AWS platform: identity, networking, LLMOps/MLOps, observability, and sovereign-hosting constraints.
- Run client workshops, write POVs and architecture packs, and stay accountable until the system is in production.
About the work
Pratham International programmes sit at the intersection of pedagogy, public systems, and production AI. The pattern is consistent across countries; the constraints change.
-
Anytime Testing Machine (ATM) — India, practice and feedback
End-to-end assessment for real classroom conditions: generate curriculum-aligned items, digitise handwritten answers (often photographed on a phone), grade against rubrics, and return feedback a teacher reviews before it reaches the learner. Built for high pupil–teacher ratios, mixed language (e.g. Hindi + English), and uneven connectivity. Pilots have already reached thousands of learners, including Second Chance (women returning for Grade 10). Model quality is measured against expert golden sets, not vendor slides. -
National assessment automation
A government-grade exam platform. Phase 1 is item authoring and print-ready paper composition before the exam window. Phase 2 is script intake, OCR at national volume, AI-assisted grading with human confirmation on the pilot, results export, and monitoring/evaluation reporting. Subjects span core and elective streams across the curriculum. Design rules that do not exist in a typical SaaS product: cycle-pinned configuration, embargoed papers, hash-chained audit trails, and integration with national student information and assessment management systems. -
Curriculum knowledge graphs
Structured retrieval corpora (concepts + worked examples) that ground generation and judging. This is the standing public-knowledge exception in an otherwise closed exam system, and the seed of a broader education DPI conversation. -
AI in TaRL and teacher support
AI tools that help teachers group and teach by learning level, with an RCT-shaped evidence bar. Same product family: human agency stays in the loop; the model is a capacity multiplier. -
Country programmes and the three-year direction
Adapt the same engine to new ministries, languages, and exam rules (Kenya and other Global South programmes are in scope). The longer bet is a learning and credentialing engine that recognises competence beyond a single syllabus.
This is not a chatbot brief. It is constrained generation, retrieval, judging, document AI, workflow, and public-sector operations.
Why this role exists
Most enterprise AI dies between the workshop and production. We need one person who owns that gap: someone who can sit with NESA, C4IR, Pratham pedagogy, and engineering in the same week, and leave behind a blueprint the delivery team can execute without inventing policy.
The target profile is a practising AI/ML Solution Architect — not a pre-sales slide owner, and not an infra-only platform engineer. The person should have done all three of these:
- Designed GenAI / NLP solutions, PoCs, and client workshops, and can talk to research-grade issues (eval, agent safety, prompt-injection, model-specific behaviour).
- Owned the AI platform: LLMOps, MLOps, DevSecOps, SRE habits, cloud migration, and FinOps.
- Taken messy discovery and turned it into Azure-centred RAG and multi-agent designs that a security team will accept.
Key Responsibilities
Discovery → architecture
- Lead solution workshops with programme, ministry, and engineering stakeholders.
- Convert half-formed requirements into service cuts, data stores, sequence diagrams, decision gates, and open-question lists with owners.
- Write architecture decision records. When the docs disagree (they will), force a decision instead of coding both.
AI system design
- Design LLM jobs as contracts: inputs, JSON schema, temperature, persist prompt hash / model / tokens / latency.
- Architect generation grounded in structured curriculum blueprints and knowledge-graph retrieval, not "ask the model for a maths paper."
- Design LLM-as-judge banks (cognitive-level, relevance, difficulty, language, fairness, standards, duplication, blueprint compliance) with thresholds, one-repair rules, and golden-dataset release gates.
- Design hybrid assembly: constraint solver first, model only on ties and layout exceptions.
- Design document AI: primary OCR provider, secondary fallback provider, confidence routing, de-identification before any grading call.
- Specify eval: inter-rater agreement, rubric-band exact match, first-pass item acceptance, cost per paper cycle.
Platform, trust, cost
- Own the reference architecture on the chosen cloud platform(s): identity, RBAC + segregation of duties, private networking, key management, object storage for scans, vector index, data warehouse for monitoring/evaluation reporting.
- Draw trust boundaries. Student PII never crosses the model gateway. Prompts are curriculum + knowledge-graph + blueprint content only.
- Put FinOps on the design: token budgets, OCR cost caps, degradation behaviour when a cap is hit, model routing.
- Specify LLMOps/MLOps: prompt/version registry, eval regression as a release gate, observability, rollback.
- Design for sovereign hosting and exam-cycle operations (embargo, watermarking, append-only audit, cycle-pinned configuration).
- Delivery and thought leadership
- Stay with the build through the first production cycle. Architecture that cannot be implemented is unfinished work.
- Produce POVs, sequence packs, and workshop artefacts that can be reused in the next country.
- Mentor tech leads on where the model should stop and the workflow should start.
Success in the first two quarters
First 30 days — Orient
- Complete architecture walkthroughs of every live/in-flight system with the current engineering leads; document current state (services, data stores, trust boundaries) as-is, not as-designed
- Meet every active stakeholder group once: engineering, programme/domain leads, and any external partner or government-side counterpart
- Inventory all open technical decisions and disagreements in existing specs; log each with current owner (if any) and status
- Identify the single highest-risk unresolved design question and flag it upward
Day 60 — Diagnose
- Produce a current-state architecture diagram covering every major pipeline: data flow, model/LLM touchpoints, storage, and identity/access boundaries
- Assign an owner and target resolution date to every open contradiction identified in month 1
- Draft a v1 reference architecture for one representative end-to-end workflow: service breakdown, data stores, identity/access model, model gateway placement, and a first-pass cost estimate
Day 90 — Commit
- v1 reference architecture is reviewed and accepted by engineering and relevant stakeholders (not just self-signed)
- Can walk any core pipeline end-to-end from memory and name every data store, every quality/eval gate, and every trust boundary in it
- At least one architecture decision record is published for a previously-contested design question, with the decision forced rather than left open
- A first cost model exists for the reference architecture, with at least one identified lever for reducing spend if usage scales
Day 120 — Harden
- Core pipeline components (generation, retrieval, judging/eval, or equivalent stages for this domain) have documented contracts: defined inputs/outputs, schema, and logged metadata (model, version, cost, latency)
- An evaluation harness exists for at least one component, with a defined pass/fail threshold
- A release gate exists that can block a bad prompt, model, or config change from reaching production
Day 150 — Extend
- The next major system phase or capability is specified to the same rigor as the first: trust boundaries, failure modes, fallback behavior, and cost caps documented before build starts
- FinOps guardrails are in place for at least one production workload: budget ceilings, degradation behavior when a cap is hit, and a routing or fallback strategy
Day 180 — Generalize - A reusable design pattern or adaptation pack exists, separating what is fixed in the architecture from what varies by deployment context (e.g. language, hosting environment, external-system integration, content/domain rules)
- At least one engineering or technical lead has been mentored on where model-based logic should stop and deterministic/workflow logic should take over
- A second production-readiness review (security, cost, or operations) has been passed on a system this person designed or substantially redesigned
Qualifications and Experience
Must have
- 8+ years building software or data/AI systems, with recent experience as a Solution Architect, AI/ML Architect, or Principal/Staff Engineer with end-to-end architecture ownership.
- Shipped at least one production AI/ML system — RAG, multi-agent, LLM-as-judge, or a comparable classical ML/NLP pipeline at scale — that passed a real security, cost, and operations review, not just a pilot or a slide deck.
- Can produce a delivery-ready architecture blueprint: services, data stores, IAM, failure modes, evaluation strategy, and cost model — detailed enough that engineering isn't guessing at intent.
- Deep cloud experience on Azure and/or AWS: comfortable defending a landing-zone design, private-endpoint architecture, and a real monthly compute/inference/OCR spend.
- Fluent in Python, with hands-on experience in both classical ML (model training, evaluation, feature pipelines) and how LLM APIs, embeddings, and vector search actually behave under production load.
- Comfortable running a room with non-engineers — domain specialists, IT stakeholders, monitoring/evaluation teams — and translating fluently in both directions.
- Strong written English; this role's output is documentation that functions as a binding technical contract.
Strongly preferred
- Azure AI (AI Foundry / OpenAI / Document Intelligence), plus one of AWS (Bedrock, SageMaker) or GCP (Vertex, BigQuery).
- LLMOps / MLOps: prompt + model versioning, MLflow or equivalent, eval harnesses, CI that can block a bad prompt.
- Document AI / OCR at volume; handwriting and low-quality scans, not only clean PDFs.
- RAG over structured knowledge (graphs, ToS, textbooks) and hybrid search, not only naive chunking.
- Education, assessment, public sector, or other high-audit domains (health, finance, identity).
- Certifications: Azure Data Scientist / Azure Solutions Architect; AWS SAA, ML Specialty, or Data Analytics Specialty.
- Experience with agent security: tool-use limits, DLP on egress, prompt-injection evaluation that is model-specific.
- FinOps discipline. You have killed a design because it would not survive the token bill.
Nice to have
- Microsoft Fabric, Databricks, Semantic Kernel / Microsoft Agent Framework, MCP (Model Context Protocol) tool integration for agentic workflows.
- Constraint solvers or exam-blueprint / item-bank systems for automated test assembly.
- Multilingual generation and feedback, including code-mixed language pairs (e.g. Hindi–English, Spanish–English, or similar).
- Experience standing up golden/eval datasets with subject-matter experts.
- Prior work with government data systems, data-sharing agreements, or in-country/sovereign hosting requirements.
What this role is not
- Not a research-only post. Publications help; shipping is the job.
- Not a pure platform/SRE hire. Infra is part of the design, not the whole design.
- Not a prompt engineer who hands off “the architecture” to someone else.
- Not a people manager of a large team on day one — you influence through artefacts and reviews.
How we will assess you
We value transparency and respect your time. Our hiring process is designed to evaluate both hands-on technical depth and architectural judgment through practical, real-world tasks.
Stage 1: Application & Knockout Screening
- Submit your CV alongside key screening responses
- Artefact — send two of: architecture note, eval design, runbook, ADR. Slides without decisions do not count.
Stage 2: Initial Conversation (30 mins)
- An introductory screen covering your background, project context and baseline alignment.
- An opportunity for you to ask questions about our technical roadmap, partner ecosystem, and vision.
Stage 3: Practical Take-Home Written Assessment (2–3 hours)
- Design exercise on a constrained generation + judge + human-review pipeline (we will use a sanitised exam-authoring brief). We want trust boundaries, eval, and cost, not a model-name dump.
Stage 4: Deep-Dive Technical & System Architecture Walkthrough
- Walkthrough of a system you shipped: what failed in review, what you changed, what it costs to run.
Stage 5: Final Stakeholder & Vision Alignment
- Final conversation with key technical leadership and program stakeholders.
Employment Details
- Function: Solutions / Architecture
- Reports to: Head of Technology / CTO
- Expected Start Date: Immediate
- Location: India (remote-first; travel to client workshops as needed)
- Experience: 8–13 years (strong 6–8 considered if you have shipped production GenAI)
- Engagement: Full-time
- Remuneration: Based on candidate’s experience and skillset
- Flexibility to work across timezones: Our team spans Latin America to Southeast Asia. Early mornings or late evenings can be a regular part of the rhythm
Why join us?
Impact at Scale: Your work will directly support the learning journeys of millions of children. With programs active in 31 countries, you will help build Digital Public Goods that bridge the gap for the millions of children across the Global South who still struggle with foundational literacy
Innovation at the intersection of Education and Technology: Lead the charge in redefining learning and learning processes through technology. You will contribute to the design and development of cutting-edge learning tools and AI-driven initiatives, ensuring EdTech isn’t just a buzzword, but a scalable solution for global equity.
A Culture of Growth and Evidence: At Pratham, we value rigorous testing and continuous learning. You will have the autonomy to experiment with the latest SOTA models, supported by a 5-year strategic roadmap and funding from partners like the Gates Foundation and Anthropic.
High-Calibre, Mission-Driven Team: You will work alongside a global team of experts from institutions like MIT, Harvard, and INSEAD. Our alumni have gone on to lead successful startups and international policy shifts, all while remaining dedicated to the goal of “Every Child in School and Learning Well.”
Inclusive and Collaborative Environment: Join a diverse team that spans Asia, Africa, and Latin America. We value innovative thinking and local expertise, ensuring that the technology we build is deeply rooted in the realities of the communities we serve.
Pratham International is an equal opportunity employer and encourages all qualified people from diverse backgrounds to apply for positions within our organisation, regardless of race, gender, disability or any other status.