Kian Heydari
Senior Developer · Machine Learning & Generative AI · Data Engineering
Senior Developer specializing in machine learning and generative AI applications, backed by solid software and data engineering: ETL pipelines, Azure migration, and CI/CD in production, plus agentic LLM pipelines and small language model distillation in research. Works across Python, SQL Server, Azure, Microsoft Fabric, Databricks, LangGraph, DSPy, and Hugging Face.
Experience
Senior IT Developer
May 2025 – Present
TD Securities · Credit Risk Reporting · Toronto, ON
- Design, maintain, and optimize ETL pipelines in Python and Microsoft SQL Server that feed credit risk reporting.
- Migrated legacy ETL applications to Azure, replacing file-watcher triggers with semantic dependencies between jobs: an easier-to-follow flow, fewer failures, and ~50% fewer jobs.
- Automate build, test, and deployment with GitHub Actions CI/CD, and keep code secure and maintainable with Snyk, Veracode, and SonarQube.
- Volunteer on the ML Tooling team, focusing on agentic AI and helping colleagues adopt the AI tools being rolled out. My team leads the organization by a wide margin on AI-integration dashboards.
- Python
- SQL Server
- Azure
- ETL
- GitHub Actions
- Snyk
- Veracode
- SonarQube
- Agentic AI
Research Assistant (Part-time)
Jan 2025 – Aug 2025
Carleton University · Supervised by Prof. Ali Arya
- Researched LLM translation for low-resource Indigenous languages such as Northern Tutchone, which mainstream models have seen little or no data for.
- Built a training-free agentic translation pipeline and a knowledge distillation pipeline for small, locally run translation models (see Projects).
- Wrote four literature reviews: LLM capabilities and efficiency; adapting generative models to low-resource languages; retrieving in-context examples without vector indexes; and grammar induction.
- Compared mT5, mBART, XLM-R adaptations (XLM-SWCM, XLM-V), NLLB, BLOOMZ, ByT5, language-specific (PhoBERT, AfriBERTa), and specialized models (FuxiMT, DeltaLM).
- LLMs
- Low-resource NLP
- Knowledge distillation
- Agentic AI
Analytics Engineer
Feb 2025 – Apr 2025
Revolution Data Platforms
- Developed a real-time customer service call management system using Python, FastAPI, and WebSockets, with end-to-end voice data pipelines.
- Prototyped voice memo transcription with structured data extraction and automated task management based on content severity.
- Optimized multi-class text classification through prompt engineering and k-fold validation.
- Improved data quality through web scraping, transformation, and preprocessing in Python.
- Python
- FastAPI
- WebSockets
- Speech-to-text
- Prompt engineering
Software Engineer
Jan 2023 – Aug 2024
Paphus Solutions
- Developed chatbot APIs across multiple platforms using Python and RESTful services.
- Optimized deep learning models for custom question answering.
- Integrated Python analytics pipelines with cloud services through optimized ETL workflows.
- Migrated a Unity-based Android application to WebGL with asynchronous connection management.
- Python
- Chatbots
- Deep learning
- REST APIs
- ETL
- Cloud
Course Instructor
Sep 2022 – Dec 2022
Carleton University
- Taught Python workshops on data manipulation, pandas, and ETL concepts.
- Guided students through ER modeling and SQL, emphasizing data integrity and query optimization.
Graduate Research Assistant
Jan 2021 – Nov 2022
Carleton University
- Built a Python data generation and transformation pipeline for deep learning research.
- Designed data preprocessing workflows and custom metrics for model evaluation.
- Python
- Deep learning
- Data pipelines
Undergraduate Research Assistant
Feb 2018 – Dec 2020
Amirkabir University of Technology
- Developed a Python anomaly detection system for time-series data, with feature extraction.
- Created simulation frameworks to generate and validate synthetic training data.
- Python
- Time series
- Anomaly detection
Education
Master of Science, Computer Science
Jan 2021 – Nov 2022
Carleton University · Specialization in Data Science
Bachelor of Science, Computer Engineering
Sep 2015 – Dec 2020
Amirkabir University of Technology
Projects
Multi-Agent Translation System
Source
Training-free, few-shot translation into languages the base LLM was not trained on, inspired by “Hire a Linguist!”. Parallel agents pull relevant dictionary entries, grammar rules, and examples; a translator agent and an LLM judge run a self-correction loop. Adds a map-reduce mode for large documents and a memory layer that learns from judge feedback.
- LangGraph
- DSPy
- Instructor
- LiteLLM
- Streamlit
Small Language Model Distillation for Translation
Teacher LLMs (Mistral, Gemma 2) enrich a small English–Northern Tutchone parallel corpus with paraphrases, context, and longer sequences; multilingual students (mT5, M2M100, mBART-50) are fine-tuned fully or with LoRA in both directions, with BLEU-based early stopping, grid search, and MLflow tracking.
- PyTorch
- Hugging Face
- LoRA
- MLflow
- Ollama
Fine-Tuning and Local RAG for LLMs
Lightweight RAG pipeline in Python for domain-specific data.
Custom Information Retrieval System
Search engine from scratch in Python, with custom indexing.
- Python
- pandas
- Information retrieval
NLP-Driven Sentiment Analysis
pandas/NumPy pipeline reaching 93.61% classification accuracy.
Skills
ML & Generative AI
- LangGraph
- DSPy
- Instructor
- LiteLLM
- Hugging Face Transformers
- PyTorch
- PEFT / LoRA
- Knowledge distillation
- RAG
- Prompt engineering
- MLflow
- scikit-learn
- pandas
- NumPy
Data Engineering
- ETL design
- SQL Server
- Microsoft Fabric
- Databricks
- PostgreSQL
- MySQL
- NoSQL
Cloud & DevOps
- Azure
- AWS
- GitHub Actions
- Docker
- Git
- Linux
- Snyk
- Veracode
- SonarQube
Backend & APIs
- FastAPI
- Flask
- Django
- ASP.NET
- REST
- WebSockets
- Streamlit
Programming
- Python
- SQL
- C# (.NET)
- Java
- JavaScript / TypeScript
Certifications & Languages
- DataCamp Data Scientist with Python
- Coursera GANs Specialization
- Languages: English (Fluent), French (Intermediate), Persian (Native)