I'm an ML Engineer with 3 years of experience in building and deploying deep learning models. I love turning data into impactful solutions, from medical image segmentation to intelligent chatbots.
Developed an intelligent RAG-based assistant for Discord that answers user questions using PDF/text documents. Integrated MiniLM + FAISS for semantic search and OpenRouter GPT for natural response generation. Enabled production-ready monitoring with Prometheus and Grafana dashboards for real-time performance tracking.
Curated 26,141 annotated samples of handwritten math expressions with Persian/Arabic numerals. Manually labeled all samples with LaTeX codes, including 6,960 complex formulas (integrals, matrices). Created comprehensive benchmark dataset for multilingual OCR research in mathematical notation.
This project implements an AWD-LSTM (ASGD Weight-Dropped LSTM) language model with 34 million parameters designed for text generation and prediction. The architecture consists of three LSTM layers with 1150 hidden units and 300-dimensional word embeddings, incorporating advanced regularization techniques including Weight Drop (applying dropout directly to recurrent weights), Locked Dropout (maintaining a consistent mask across all time steps), and Embedding Dropout. Additionally, weight tying between the embedding and output layers is employed to reduce the total number of parameters while maintaining model performance. The model was trained on the WikiText-103 dataset, which contains over 1.8 million training samples extracted from more than 1,800 Wikipedia articles. The training process utilizes an SGD optimizer with an initial learning rate of 30, momentum of 0.9, and weight decay of 1.2e-6. Gradient clipping at 0.25 and ReduceLROnPlateau learning rate scheduling ensure stable convergence throughout training. The final model achieves a validation perplexity of approximately 108, demonstrating strong language understanding capabilities with the ability to generate fluent, grammatically correct, and contextually coherent text completions.
This Image Captioning model uses an Encoder-Decoder architecture. The encoder is a pre-trained ResNet50 that extracts visual features from input images into 2048-dimensional vectors. The decoder is a 2-layer LSTM with 512 hidden units that generates captions word by word, using an embedding layer of size 256. Dropout of 0.5 is applied to prevent overfitting. The model was trained on the Flickr30k dataset containing 31,000 images with 5 captions each. The training used Cross-Entropy Loss with AdamW optimizer (lr=0.0001) for 20 epochs with batch size 64. Teacher forcing was applied with 0.8 probability. The final model achieved a validation loss of 2.34 and BLEU-4 score of 0.31. The trained model is containerized with Docker and deployed as a web app on Hugging Face Spaces.
A deep learning-based web application that generates descriptive captions for any uploaded image. Built with Flask backend, ResNet50 encoder + LSTM decoder architecture, trained on Flickr30k dataset. Containerized with Docker and deployed on Hugging Face Spaces.
An AI-powered autocomplete web application that predicts the next word(s) in real-time as you type. Built with Flask backend, AWD-LSTM language model architecture with 34 million parameters, trained on WikiText-103 dataset (1800+ Wikipedia articles). Features dynamic phrase suggestions (up to 4 words), adjustable creativity (temperature) control, and real-time prediction triggered by space key. Deployed with Docker on Hugging Face Spaces.