
Sakib Samar
Software Engineer
By day at HPE, I build and scale the backend systems and high performance storage that keep AI and research workloads fed at supercomputing scale. By night, I tinker on my own projects.
About
I'm a Software Engineer at Hewlett Packard Enterprise, building distributed storage systems for large-scale AI and HPC environments. My work spans control planes, backend services, and storage infrastructure, with a focus on reliability, performance, and operating at scale.
Areas of focus
- Go
- Python
- Java
- SQL
- REST APIs
- Authentication
- RBAC
- Async Orchestration
- etcd
- Concurrency Control
- Distributed Consistency
- Fault-Tolerant Workflows
- Lustre
- DAOS
- Kubernetes
- Docker
- Linux
- Redfish
- Swordfish
- Prometheus
- Grafana
- OpenSearch
Experience
Where I've worked

June 2022 – Present
Software Engineer
Building the control plane and storage performance stack behind HPE's exascale AI and HPC systems — from Go services that manage 10,000+ node clusters to the I/O path that feeds GPU training jobs.
- Built the Go REST API layer for a multi-tenant storage control plane (DMTF Redfish, SNIA Swordfish) managing DAOS clusters at 10,000+ node scale, with auth, role-based access, and etcd-backed state.
- Led HPE's first two MLPerf Storage submissions on Lustre — the latest sustaining 160 GB/s across 96 accelerators — and helped shape how the industry benchmarks storage for AI training.
- Shipped Lustre hybrid I/O in the ClusterStor release: dynamic buffered-to-direct promotion with tunable size thresholds and submission concurrency, lifting sequential read throughput for training data loading.
- Designed an event aggregation and alerting pipeline spanning every DAOS service and node (Grafana, OpenSearch, HPCM), cutting incident response time by roughly 90%.
- Made hardware-dependent workflows safe to change: fault-tolerant, multi-step orchestration with compensating rollback, plus injectable command interfaces that unit test without physical clusters.

February 2022 – June 2022
Software Engineer Co-Op
Backend work on smartcard identity and authentication services.
- Built a Java smartcard profile token tool that extended certificate lifetime and authentication for existing users.
- Cut webservice build time 22% by reworking build dependencies and artifacts.
- Migrated logging from Log4j to Logback, closing a critical security vulnerability, and tightened input validation to reduce auth latency.

November 2021 – January 2022
Software Development Intern
Server-side development for corporate Android client apps.
- Built server-side services backing corporate client applications running on Android devices.
- Shipped a recommendation service that surfaced related offerings from customer profiles.

October 2021 – May 2022
Machine Learning Research Assistant
Computational modeling and ML for parallel optical processing.
- Implemented the Laguerre-Gauss algorithm in Python to model orbital angular momentum (OAM) states.
- Applied machine learning to simulation data to analyze parallel optical processing.
Talks & Publications
Sharing lessons learned from benchmarking, distributed storage, and AI infrastructure.

Keynote
AI/ML Benchmarking and Lustre
Lustre User Group 2024
Presented HPE's approach to benchmarking AI and ML workloads on Lustre-based storage systems, covering methodology, performance considerations, and lessons learned.
Watch Presentation →
Publication
E2000 Performance: From Microbenchmarks to Applications
Cray User Group 2025
Explores storage performance characteristics across synthetic benchmarks and real-world application workloads, highlighting practical observations from large-scale HPC environments.
Read Paper →