Portrait of Salman Farcy Peer

Hi, I'm Salman Farcy

Data Scientist

I build AI for finance that people can actually defend. Three years building production credit risk models, now applying the same discipline to retrieval and agents.

About Me

I’m a Data Scientist with 3+ years of experience building, deploying, and monitoring production machine learning models for credit and fraud risk in retail consumer lending. I’ve built acquisition and Weight of Evidence (WOE) scorecard models that score loan applications in production, engineered features from complex data sources like credit bureau and bank aggregator data at scale, and partnered with cross functional teams to move models from development into a monitored production environment. I work within regulatory guidelines, grounded in reproducible research and model governance. I’m currently doing my Master’s in Data Science at the University of Maryland, working at the intersection of applied AI and finance with production RAG systems, LLM evaluation, and agentic workflows across credit risk, fraud, and AML, to build trustworthy AI.

  • Python
  • SQL
  • XGBoost
  • LightGBM
  • Spark
  • Databricks
  • LangGraph
  • RAG
  • AWS

Featured Projects

Each one carries a measured result from a held out evaluation, and a case study explaining the decisions behind it.

  • Python
  • LangGraph
  • XGBoost
  • NetworkX
  • OpenAI API
  • Pydantic

Agentic fraud and AML alert review that measures the recall cost of every case it closes.

Read more
  • Python
  • RAG
  • ChromaDB
  • SQLite
  • Hybrid search
  • Cross-encoder reranking

Dual engine hybrid RAG over SEC 10-K filings. Grounded, cited, and it refuses when the filings do not support an answer.

Read more
  • Python
  • Feature engineering
  • Model stacking
  • Hypothesis testing
  • scikit-learn
  • LightGBM

A credit default pipeline built the way a lender actually has to defend it.

Read more

Experience & Education

  1. Sep 2022 — Jun 2025

    Data Scientist, Credit Risk at HCLTech

    1. Built and deployed three consumer credit acquisition models (XGBoost, LightGBM, logistic regression) into production through the ACTICO Business Rule Engine with H2O and PMML, improving the lead to disbursal ratio by 9 percent.
    2. Developed a prospect initial credit limit assignment model, merging Balance and ADB models into a unified multi target regression with embedded causal inference, reducing model inventory, execution latency and production infrastructure cost by 25 percent.
    3. Engineered 1,700+ features across roughly 100K loan applications by unifying five data sources, credit bureau, bank aggregator, web analytics, SMS and email, with Weight of Evidence encoding and Spark pipelines on Databricks, cutting feature build time by 25 percent.
    4. Conducted comprehensive model evaluations using decile level accuracy, Gini coefficient and SHAP based feature interpretability to assess performance, identify key drivers and enhance model transparency.
    • XGBoost
    • LightGBM
    • Spark
    • Databricks
    • PMML
    • WOE encoding
  2. Feb 2026 — Present

    Graduate Assistant at University of Maryland

    1. Built the retrieval layer of a grounded, FERPA compliant RAG assistant over 5,000+ de-identified counseling case documents, using semantic chunking, vector embeddings and metadata filtering by clinic, year and document type so queries resolve to the exact relevant records instead of scanning the full corpus.
    2. Engineered grounded answer generation that draws only from retrieved sources, cites every claim and refuses when the records do not support an answer, eliminating hallucinated responses in a privacy sensitive clinical setting.
    3. Deployed the system across 40 to 50 clinics, cutting analyst review time from hours to seconds.
    • RAG
    • Vector search
    • Semantic chunking
    • Python
    • FERPA compliance

Technical Skills

ML and Modeling

  • Python
  • SQL
  • XGBoost
  • LightGBM
  • scikit-learn
  • PyTorch
  • TensorFlow

Applied GenAI

  • RAG
  • LangGraph
  • LangChain
  • ChromaDB
  • LLM evaluation
  • Agents
  • Hugging Face

Domain

  • Credit risk
  • Fraud
  • AML
  • Financial NLP
  • Anomaly detection
  • Calibration
  • Model validation

Data and Cloud

  • Spark
  • Databricks
  • AWS
  • FastAPI
  • Docker
  • Git

Get In Touch

Open to Data Scientist and AI Engineer roles. The fastest way to reach me is the email below.