Cristian GHITU
Data Scientist | Machine Learning Engineer
Data Scientist and Machine Learning Engineer with over 6 years of experience in the banking sector,
specializing in the end-to-end delivery of Big Data and ML applications.
Proven ability to bridge the gap between business stakeholders and technical execution,
from scoping requirements and building predictive models to refactoring code into modular,
production-ready data pipelines.
Expertise in Python, Scala and Apache Spark in a Big Data environment,
with a strong focus on industrialization.
Do I Match ?
Are you a recruiter? I built this app to help you check if my profile matches your job description! Try it out!
Professional Experience
Course Creator, Personal Projects & Career Break
GenAI and LLMs: built an app that checks if a provided job description and a user’s resume match,
providing a summary and a list of bullet points justifying the decision.
The project involved using LangChain as a generic interface for models
(tested with OpenAI-provided models) with structured output. FastAPI is used for programmatic requests,
with a Streamlit app provided as a more user-friendly interface.
A similar project involved creating a pipeline that checked several platforms for job
descriptions matching a given candidate, with the possibility of checking company details for
successful matches. Main technologies and services involved: Python, Selenium, LangChain,
OpenAI, Tavily.
Course Development: During a professional sabbatical, designed and built a comprehensive online introductory course to the Scala programming language, featuring video lectures, slides, interactive exercises and quizzes.
Technical Implementation: Built Scala unit tests to automatically evaluate student exercise submissions for correctness and utilized the host platform’s interface to implement auto-evaluated quizzes.
Upskilling: Achieved the AWS Cloud Practitioner certification. Upskilled in ML Engineering and GenAI tools, including Docker (containerization), GitHub Actions (CI/CD), Apache Airflow (orchestration), FastAPI (model serving and APIs), Streamlit (web applications), PyTorch, Langchain, LangGraph (LLM and Agentic AI).
Data Scientist
Agile Framework:
Executed Data Science projects within Agile teams using the Scrum methodology,
driving initiatives from the scoping phase through to industrialization.
Key Projects:
Risk Scoring:
Built industrialized ML pipelines to calculate risk scores.
Conducted feature engineering on large datasets with PySpark and trained models using scikit-learn,
Hyperopt and other ML libraries, leveraging Shapley values for explainability.
Structured data processing code into modular pipelines using Kedro,
with Great Expectations used for quality checks.
Used MLflow for versioning models and tracking performance across different iterations.
Transaction Processing Automation:
Developed a scoring model to automate banking transactions processing.
Translated Python code into Scala for productionizing the solution.
Regulatory KPIs:
Extracted and crossed massive volumes of tabular and textual data
to estimate complex regulatory indicators using PySpark.
Billing Anomaly Detection:
Applied unsupervised anomaly detection models to billing datasets
to identify financial irregularities.
Incident Classification via NLP (Early PoC):
Applied natural language processing to group and classify IT incidents.
Used Apache Spark and Scala with OpenNLP and StanfordNLP libraries
to leverage distributed processing and reduce execution time for textual preprocessing.
Engineering Culture & Tooling:
Workflow Standardization:
Standardized ML workflows to accelerate the transition
from Jupyter Notebook PoCs to industrialized applications,
establishing robust conventions for exploratory code.
Mentorship & Community:
Fostered knowledge sharing as an active member of data communities
by organizing internal and open meetups,
as well as mentoring junior data scientists during onboarding.
Course Instructor: Scala and Apache Spark
Delivered an introductory Scala and Apache Spark course
to an audience of over 300 Data Science and Engineering students.
Taught in both English and French.
Data Science Intern
Performed anomaly detection on indicators extracted from application logs
to enable predictive server maintenance.
Explored unsupervised methods including Taylor double seasonality models,
LSTMs, and auto-encoders using Python and Keras.
Education & Certifications
Engineering Degree in Computer Science
University of Technology of Compiègne (UTC), France • 2012 - 2017
AWS Certified Cloud Practitioner
Amazon Web Services • Issued 2025
Skills
Add me on LinkedIn!
Let's exchange ideas, talk about collaboration opportunities or simply stay in touch!