Job Description – AI Engineer Intern (Agentic AI)


Role: AI Engineer Intern – Agentic AI & LLM Engineering
Department: AI Engineering
Location: Bengaluru
Employment Type: Internship

About the Role

We are looking for a technically strong and passionate AI Engineer Intern with hands-on experience building Agentic AI applications using Large Language Models (LLMs).
The ideal candidate should have practical exposure to DeepAgents, LangGraph, Agent Harness development, Model Context Protocol (MCP), AI Observability, and LLM Evaluations (Evals) through academic projects, personal projects, internships, or open-source contributions.
You will work on designing, developing, evaluating, and deploying AI agents capable of performing complex, multi-step business workflows by interacting with tools, APIs, databases, and enterprise systems.
We are looking for candidates who go beyond basic chatbot development and are interested in building reliable, tool-enabled, and measurable AI applications.

Key Responsibilities

1. Agentic AI & Agent Harness Development

  • Design and develop AI agents using LangGraph, LangChain, and DeepAgents.
  • Build and customize agent harnesses to support planning, tool execution, context management, and multi-step workflows.
  • Implement agent middleware for tool retries, error handling, human-in-the-loop approvals, and context management.
  • Work with agent memory, state management, checkpointing, and persistent execution.
  • Develop reusable tools, sub-agents, and agent orchestration workflows.

2. Model Context Protocol (MCP) & Tool Integration

  • Develop and integrate MCP servers and clients to connect AI agents with external systems.
  • Build reusable tools for interacting with APIs, databases, and enterprise applications.
  • Implement structured tool calling, input validation, error handling, and secure tool execution.
  • Understand tool discovery, tool permissions, and integration patterns.

3. AI Observability & Monitoring

  • Implement end-to-end tracing and monitoring for LLM and agent-based applications.
  • Work with observability frameworks such as MLflow, LangSmith, Langfuse, or OpenTelemetry.
  • Capture and analyze agent execution traces, tool calls, model responses, token consumption, latency, and failures.
  • Debug agent execution paths and identify reliability and performance issues.
  • Support monitoring and optimization of AI application quality and inference costs.

4. LLM Evaluations (Evals) & Reliability

  • Design and implement evaluation frameworks for AI agents and LLM applications.
  • Create evaluation datasets with ground-truth examples and expected outputs.
  • Implement evaluations for tool selection, tool execution, structured outputs, and end-to-end agent decisions.
  • Measure correctness, precision, recall, consistency, hallucinations, and task completion.
  • Develop automated regression tests to evaluate changes in prompts, models, and agent workflows.
  • Experiment with LLM-as-a-Judge and human feedback mechanisms.

5. Generative AI & Application Development

  • Develop LLM-powered applications using Python and model APIs.
  • Implement prompt engineering, structured outputs, function calling, and Retrieval-Augmented Generation (RAG).
  • Build APIs and backend services using FastAPI or similar frameworks.
  • Work with structured and unstructured enterprise data, including documents, emails, PDFs, and databases.
  • Collaborate with engineering teams to develop and deploy AI solutions for real-world business use cases.

Required Technical Skills

  • Programming: Strong Python fundamentals, OOP, asynchronous programming, REST APIs, and JSON.
  • Agentic AI: Hands-on exposure to LangGraph, DeepAgents, or comparable agent frameworks, including agent orchestration and tool calling.
  • Agent Harness: Understanding of agent execution loops, middleware, tool integration, state management, and error handling.
  • MCP: Understanding of MCP architecture and experience building or integrating MCP servers and tools.
  • Observability: Familiarity with instrumenting and debugging LLM or agent workflows using tracing and monitoring tools.
  • Evals: Understanding of evaluation datasets, correctness metrics, and methods for evaluating LLM and agent outputs.
  • LLMs: Knowledge of prompts, embeddings, structured outputs, function calling, and RAG.
  • Software Engineering: Git, debugging, unit testing, and basic database knowledge.

Good to Have

  • Experience with Azure OpenAI, OpenAI APIs, Anthropic, or open-source LLMs.
  • Familiarity with MLflow, LangSmith, Langfuse, or OpenTelemetry.
  • Experience with Pydantic, FastAPI, PostgreSQL, MongoDB, or vector databases.
  • Knowledge of Docker, cloud deployment, and CI/CD pipelines.
  • Exposure to multi-agent architectures, human-in-the-loop workflows, and persistent agent memory.
  • Familiarity with document intelligence, information extraction, or enterprise workflow automation.

Educational Qualifications

  • Currently pursuing or recently completed B.E./B.Tech/M.Tech/MCA in Computer Science, Artificial Intelligence, Data Science, Information Technology, or a related discipline.
  • Strong programming fundamentals and demonstrated interest in AI engineering.
  • Academic projects, open-source contributions, hackathons, or independent AI projects are an advantage.

Ideal Candidate

We are looking for someone who:
  • Has explored building AI agents beyond basic prompt-based applications.
  • Understands how agents interact with tools, manage state, and execute workflows.
  • Thinks about AI reliability, evaluations, and observability.
  • Can debug agent workflows and identify opportunities for improvement.
  • Enjoys working on real-world engineering challenges.
  • Is proactive, technically curious, and comfortable exploring rapidly evolving AI technologies.
  • Can explain the architecture, implementation, and technical decisions behind their projects.
  • Our ideal candidate is a hands-on AI builder who can design an agent, connect it to real tools, evaluate its decisions, observe its execution, and continuously improve its reliability.