Submitting more applications increases your chances of landing a job.

Here’s how busy the average job seeker was last month:

Opportunities viewed

Applications submitted

Keep exploring and applying to maximize your chances!

Looking for employers with a proven track record of hiring women?

Click here to explore opportunities now!
We Value Your Feedback

You are invited to participate in a survey designed to help researchers understand how best to match workers to the types of jobs they are searching for

Would You Be Likely to Participate?

If selected, we will contact you via email with further instructions and details about your participation.

You will receive a $7 payout for answering the survey.


User unblocked successfully
Thank you. Your report has been submitted and will be reviewed shortly.
Harish  Janagam , Senior Agentic AI Engineer | Full-Stack Builder

Harish Janagam

Senior Agentic AI Engineer | Full-Stack Builder·Agentic Contract Intelligence Platform

India

Bachelor's degree, Electronics & Communication Engineering

Work experience

Total years of experience: 7 years, 9 months

Senior Agentic AI Engineer | Full-Stack Builder

January 2025 - Present

Agentic Contract Intelligence Platform

Bengaluru, India

January 2025 - Present

Project Summary: Full-stack Agentic AI platform targeting legal, procurement, and delivery teams handling IT contracts (SOW, SOWA,
SLA, MSA). Designed for production-grade reliability with enterprise-level observability and zero-trust document integrity.
Tech Stack: React + TypeScript + Vite
• FastAPI
◦ CrewAI
◦ Docling
◦ PostgreSQL (asyncpg + pgvector)
◦ Qdrant
◦ OpenAI GPT-4o
◦ Celery
◦ Alembic
• Docker
▸ Architected a CrewAI-based hybrid multi-agent orchestration layer with specialist agents for clause extraction, risk scoring,
obligation mapping, and anomaly detection
▸ Designed a 4-agent crew (Contract Analyst, Risk Assessor, Obligation Mapper, Report Generator) with inter-agent task delegation
and structured handoff protocols
▸ Built a dual-store retrieval architecture linking PostgreSQL clause UUIDs with Qdrant HNSW vector indexes — enabling sub
second semantic search with no extra DB roundtrips
▸ Implemented a custom ContractID schema (YYYYMMDD_HHMMSS-) with 100% verbatim clause storage for
audit integrity and tamper detection
▸ Parallelised OpenAI extraction across chunked clause batches to process 500-page contracts in under 60 seconds end-to-end
▸ Replaced unstructured.io with Docling for local document parsing (zero cloud cost, zero data egress) supporting PDF, DOCX, and
multi-column layouts
▸ Resolved enterprise TLS inspection failures by injecting OS certificate store via the truststore library; suppressed CrewAI OTEL
telemetry noise in corporate environments
▸ Designed Azure cloud deployment architecture (Azure Container Apps, Azure Document Intelligence, Azure Cosmos DB) with
detailed cost modelling across three deployment tiers
▸ Built React + TypeScript + Vite frontend with contract upload, multi-tab clause explorer, risk dashboard, and obligation timeline
views Alight (Healthcare Analytics)

Company industry:
IT Services

Senior Data Scientist

January 2024 - Present

Optum–UHG

Bengaluru, India

January 2024 - Present

Project Summary: End-to-end Agentic Contract Analysis Platform for a leading US healthcare PBM. Enables legal teams to upload,
compare, and validate contract documents with AI-driven term extraction and structured feedback workflows.
Tech Stack: Python
• FastAPI
◦ LangGraph
◦ LangChain
◦ Azure OpenAI (GPT-4o/4o-mini)
◦ Azure Document Intelligence
◦ Azure AI Search

MongoDB
• GraphQL
◦ Pydantic
▸ Designed and deployed a multi-agent LangGraph pipeline with dedicated agents for contract term extraction, email notification,
and feedback control
▸ Built specialised tool functions integrated with agents for clause-level extraction across PBM pricing tables and contractual
definitions
▸ Orchestrated stateful LangGraph nodes with conditional edges for dynamic routing between extraction, validation, and user
notification phases
▸ Achieved 100% accuracy on term and pricing table extraction by combining Azure Document Intelligence layout analysis with
GPT-4o structured outputs
▸ Implemented an email notification sub-agent within the graph to deliver real-time alerts to content writers and auditors upon
document processing
▸ Designed human-in-the-loop feedback node allowing legal reviewers to flag, correct, and re-inject validated terms back into the
pipeline
▸ Applied Azure AI Search for vector embeddings and similarity retrieval across large contract corpora; MongoDB for persistent
clause and audit storage

Company industry:
IT Services

Data Scientist

January 2021 - January 2024

Cignex Datamatics

Bengaluru, India

January 2021 - January 2024

Project Summary: Predictive modelling on patient diagnosis data for 50+ major US healthcare clients including UHG, BCBS, Cigna, and
Athena. Built hyper-personalisation and classification models on insurance, claims, and 401K datasets at scale using Databricks and
Spark.
Tech Stack: Python
• Databricks
◦ Apache Spark
◦ PySpark
◦ SparkSQL
◦ MLlib
◦ PyTorch
◦ TensorFlow
◦ AWS SageMaker
◦ S3
◦ Hue/Impala

LangChain
• HuggingFace Falcon
▸ Engineered large-scale data pipelines on Databricks using PySpark and SparkSQL to process and merge multi-source healthcare
datasets (claims, diagnosis, 401K) across distributed Hue/Impala databases and S3
▸ Performed complex Spark transformations — joins, window functions, aggregations, and partitioning — on 50+ client datasets to
produce hyper-personalised feature matrices for downstream ML models
▸ Optimised Spark job performance through broadcast joins, caching strategies, and partition tuning, reducing pipeline runtimes
significantly on large patient cohort data
▸ Built and deployed PyTorch and TensorFlow deep learning models for patient classification and risk stratification on structured
clinical and insurance data
▸ Developed and deployed multiple classification and regression models (XGBoost, LightGBM, Random Forest) on patient
diagnosis, claims, and new-hire insurance datasets
▸ Built a GenAI Q&A pipeline using LangChain over PDF clinical data; created a HuggingFace Falcon-powered text generation model
for internal knowledge retrieval
▸ Implemented NER-NLP pipelines for healthcare text classification, extracting structured entities from unstructured clinical notes
▸ Managed end-to-end ML lifecycle on AWS SageMaker: data wrangling, feature engineering, model training, deployment, and
monitoring across 50+ healthcare client environments

Company industry:
IT Services

Data Scientist

December 2020 - January 2021

Future Focus Infotech

Bengaluru, India Remote

December 2020 - January 2021

▸ Created supply-and-demand ML models for BmaaS, MonaaS, DBaaS, and CaaS service pricing prediction
▸ Dockerised microservices and established pipeline architecture following MVC patern for multi-service orchestration
▸ Migrated aiohtp-based services to Flask framework; built RESTful APIs to ingest real-time service telemetry (Financial

Company industry:
IT Services

Data Analyst

January 2020 - January 2020

Atyeti Inc.

Hyderabad, India Remote

January 2020 - January 2020

Stack: Python
• Pandas
◦ ETL
◦ SSMS
◦ Jupyter
▸ Executed content migration from legacy to updated financial data services with automated data comparison and validation
pipelines
▸ Extracted and transformed structured financial datasets for production deployment across FactSet content systems

Company industry:
IT Services

Consultant – Data Science

January 2019 - January 2020

GyanSys Inc.

Bengaluru, India Hybrid

January 2019 - January 2020

Python Pytesseract
• Scikit-learn
◦ Linear Regression
◦ Pandas
◦ Jupyter
▸ Developed regression models to predict soap bar hardness across chemical compositions and temperature conditions, providing
R&D insights for multiple product brands
▸ Analysed how different chemical parameters react under varying weather conditions to model soap bar behaviour at scale
▸ Automated data extraction and consolidation from multi-sheet Excel research files, cleaning and aligning data to business
requirements
▸ Built an invoice image parsing pipeline using Pytesseract to extract structured fields (quantities, values, dates) from scanned
invoice images
▸ Applied Linear Regression and statistical methods for outlier detection, missing value imputation, and predictive quality scoring
on R&D datasets

Company industry:
Other Business Support Services

Process Lead / Data Science Associate

January 2017 - January 2018

Infosys

Bengaluru, India Hybrid

January 2017 - January 2018

• Scikit-learn
◦ Power
◦ SQL
▸ Analysed hotel and resort booking data to uncover customer behaviour paterns and provide business intelligence to account
managers
▸ Developed regression models for customer booking rate prediction; visualised insights in Power BI dashboards for executive
presentations
▸ Collaborated cross-functionally with data science and business stakeholders within a banking application (Finacle) context

Company industry:
IT Services

Education

Jawaharlal Nehru University

January 2013

January 2013

Bachelor's degree, Electronics & Communication Engineering

India

Skills

ARTIFICIAL INTELLIGENCE

Intermediate

COMPUTER SYSTEMS

Intermediate

CONSULTING

Intermediate

DATA SCIENCE

Intermediate

END TO END ENCRYPTION

Intermediate

HEALTHCARE SERVICES

Intermediate

LANGUAGE MODEL

Intermediate

LEADERSHIP

Intermediate

MEDICATION PROMPTING

Intermediate

MLOPS MACHINE LEARNING OPERATIONS

Intermediate

Languages

English

Beginner

Telugu

Beginner

Hindi

Beginner

Training and Certifications

Certifications
Azure Data Scientist Associate (DP-100)