Featured Project

Scientific Consulting Portfolio

Real consulting deliverables — an exploratory data analysis report and a full machine learning pipeline — showing exactly what a client receives when they work with STEM-azing Scientific Consulting. Every analysis below was designed, coded, and written personally by Dr. Delva.

STEM-azing Scientific Consulting

The independent practice of Nella C. Delva, PhD — a Fulbright Research Fellow and biomedical scientist based in Berlin, Germany. The practice specializes in the intersection of rigorous life science expertise and applied data analytics, offering consulting, analysis, and writing services to academic researchers, clinical teams, and biotech organizations worldwide. Every deliverable is produced personally by Dr. Delva — not delegated, not templated.

PhD
Biomedical Sciences /
Neuroscience, FSU
Fulbright
Research Fellow,
MDC Berlin (2025–26)
McKnight
Merit-based national
doctoral fellowship
Le Wagon
Data Science & AI
Bootcamp, Berlin
3
Languages —
EN / FR / Haitian Creole

Six areas of practice, three engagement tiers

Full descriptions and pricing live on the Services page — here's the shorthand.

📈 Data Analysis & Visualization

EDA, statistical testing, publication-quality figures, Python/R pipelines.

🤖 Machine Learning & Predictive Modeling

Classification & regression, feature engineering, model validation, biological ML pipelines.

✍️ Scientific Writing & Grant Support

Manuscript development, grant proposals (DFG, ERC, NIH), research reports.

🧠 Research Consulting & Study Design

Experimental design strategy, hypothesis development, methods selection.

🎓 Academic Mentoring

Doctoral & postdoc coaching, fellowship application prep, career transition support.

📋 Reporting & Dashboards

BI dashboards, GA4/GTM analytics reporting, SQL/BigQuery pipelines, data storytelling.

Three levels of consulting depth

Starter

Entry-level engagement

  • EDA report (PDF + Word)
  • 3 publication-quality figures
  • Python/Jupyter notebook
  • Written findings & next steps
★★

Standard

Mid-tier engagement

  • Everything in Starter
  • Multivariable modeling
  • Subgroup & interaction analyses
  • 5+ figures + data tables
  • 1 round of revisions

Exploratory Data Analysis Report

Sleep Quality & Stress Biomarker Dataset · N = 180 Participants

6.8 hrs
Mean Sleep Duration
5.2 / 10
Mean Stress Score
90 / 90
Control / Intervention (N)
r = −0.43
Sleep–Stress Correlation
38%
Below Recommended Sleep

Project Overview

An exploratory analysis of a participant health dataset examining the relationship between sleep duration, stress levels, and intervention group assignment. The client required a rigorous statistical summary, publication-quality visualizations, and an interpretation grounded in current sleep science literature.

Research Questions

  • Is sleep duration significantly associated with stress levels in this population?
  • Does the Intervention group show meaningfully different sleep patterns vs. Controls?
  • What proportion of participants fall below recommended sleep thresholds?

Engagement Details

Service Tier★ Starter
Dataset SizeN = 180 participants
VariablesSleep duration, Stress score, Group
Analysis TypeExploratory / Descriptive
LanguagePython (Jupyter)
DeliverablesReport + Code + 3 Figures

Tools & Methods

  • Python (pandas, NumPy, SciPy, Matplotlib)
  • Pearson correlation, two-tailed significance testing
  • Independent samples t-test; Cohen's d effect size

Descriptive Statistics

VariableValuep-valueInterpretation
Mean Sleep Duration6.8 hrsBelow NSF recommended 7–9 hrs for adults
Mean Stress Score5.2 / 10Moderate stress level across full sample
Sleep–Stress rr = −0.43< 0.001Significant negative correlation
Intervention vs. Control Sleep+0.8 hrs< 0.05Intervention group shows modestly higher sleep
Cohen's d (group effect)d = 0.38Small to moderate effect size
% < 6 hrs sleep/night38%Below minimum recommended threshold

Key Findings

1
Sleep duration was negatively correlated with stress score (r = −0.43, p < 0.001), indicating that higher stress is robustly associated with reduced sleep across the full sample — significant even after accounting for group assignment.
2
The Intervention group achieved modestly higher sleep duration (+0.8 hrs) vs. Controls. Statistically significant, but the effect size was small (Cohen's d ≈ 0.38) — a real but not transformative improvement.
3
Both groups showed comparable variance in sleep patterns — the intervention's benefit was not universal, pointing toward subgroup analyses as a valuable next step.
4
Approximately 38% of the full sample reported fewer than 6 hours of sleep per night — a clinically significant prevalence representing a priority subgroup for targeted follow-up.

Scientific Interpretation

The data provides clear evidence of a stress-sleep relationship in this sample. The intervention appears to improve sleep duration modestly, but the small effect size and heterogeneous response suggest that not all participants are equally responsive as designed. The 38% prevalence of severely restricted sleep indicates a meaningful public health burden within this cohort.

Recommended Next Steps

  • Collect follow-up data at 3 and 6 months to assess durability
  • Add validated sleep quality measures (PSQI) alongside duration
  • Control for age, gender, and baseline stress in multivariable models
  • Conduct subgroup analyses to identify highest-responding profiles
  • Expand sample to N ≥ 400 to improve statistical power
Written Report
(PDF + Word)
Jupyter Notebook
(Reproducible)
3
Publication-Quality
Figures (300 DPI)
5 days
Turnaround

ML Predictive Analysis Report

Sleep Disorder Risk Classification · N = 300 Participants · 7-Feature Biomarker Dataset

Project Scope & Objectives

Building a full machine learning pipeline to predict sleep disorder risk from a multi-modal biomarker and lifestyle dataset. Three classification models were developed, compared, and validated. Feature importance and partial dependence plots connected statistical findings back to the underlying neuroscience and physiology.

Research Questions

  • Which biomarkers and lifestyle factors best predict high sleep disorder risk?
  • How do three ML models compare on held-out data?
  • What is the non-linear relationship between sleep duration and predicted risk?
  • Can a lightweight screening tool be derived for clinical use?

Engagement Details

Service Tier★★★ Advanced
Dataset SizeN = 300 participants
FeaturesSleep, Stress, Age, Cortisol, BDNF, CRP, IL-6
TargetSleep Disorder Risk (High vs. Low)
Models Built3 (Logistic Regression, Gradient Boosting, Random Forest)
Validation5-fold CV + held-out test set
DeliverablesReport + Code + 6 Figures + CSVs + 2 Revisions

Model Performance Summary

ModelAUC-ROCAccuracyCV ScoreRecommendation
Logistic Regression (baseline)0.79172.3%0.784 ± 0.04Good baseline; interpretable coefficients
Gradient Boosting0.83476.7%0.826 ± 0.03Strong performance; handles non-linearity
★ Random Forest (selected)0.84778.0%0.839 ± 0.03Best performer; selected as final model

Feature Importance & Biological Interpretation

Sleep Duration
0.28
Stress Score
0.22
Cortisol
0.18
BDNF
0.14
CRP
0.10
IL-6
0.05
Age
0.03

Sleep duration is the primary protective factor — chronic restriction is a leading risk factor for sleep disorders. Stress score reflects HPA axis activation disrupting sleep architecture via cortisol elevation; cortisol, BDNF, CRP, and IL-6 add independent biological signal linking inflammation and neuroprotection to sleep outcomes.

Key Findings & Decision-Support Recommendations

🎯
The Random Forest model achieves AUC = 0.847 with 78% accuracy on held-out data — strong performance for a 7-feature biological dataset without deep learning.
💤
Sleep duration and stress score together account for 50% of the model's discriminative power — both modifiable, high-priority intervention targets.
📉
The partial dependence plot reveals a critical non-linear effect: below 6 hours of sleep, predicted risk increases sharply. The 6–7 hour window is the highest-leverage intervention zone.
🧬
Cortisol and BDNF add independent predictive value beyond sleep and stress alone, suggesting targeted biomarker screening could meaningfully improve clinical risk stratification.
A composite score of Sleep Duration + Stress Score + Cortisol captures 68% of the full model's predictive power — deployable as a lightweight clinical screening tool without full biomarker panels.
Full ML Report
(Word + PDF)
Python Jupyter
Notebook
6
Publication-Quality
Figures
Model Eval
Spreadsheet + CSVs
2
Rounds of
Revisions

Every project receives full scientific rigor, clear communication, and timely delivery

Whether you need a rapid exploratory report or a full predictive modeling pipeline, engagements are scoped to match your data, your timeline, and your budget.

See pricing & book a service → Book a 20-min call →