About Anton
Hi, I’m Anton Siomchen—a medicinal chemist and software engineer working at the intersection of chemistry, data, and software development.
I’m currently a Senior Software Engineer specialising in chemoinformatics at EPAM Systems. My work focuses on building practical software for chemical data, virtual screening, chemical-space exploration, and computer-aided drug discovery.
What I work on
I’m particularly interested in:
- Cheminformatics software architecture
- RDKit and Python tooling
- Chemical databases and PostgreSQL cartridges
- Molecular fingerprints, descriptors, and similarity search
- Reproducible drug-discovery workflows
- Scientific software usability and documentation
Open-source projects
A few projects and contributions worth highlighting:
- MolAlchemy — a Python library connecting SQLAlchemy with RDKit and Bingo chemical database cartridges. It provides typed molecular and reaction fields alongside chemical queries and database tooling.
- Mols2Bases — an Obsidian plugin for importing molecular datasets, rendering structures with RDKit.js, and filtering them using text or SMARTS queries.
- RDKit — contributions to fixes and improvements in the open-source cheminformatics toolkit.
- Scikit-Mol — documentation and developer-experience improvements for its RDKit and scikit-learn integration.
Background
I hold an MSc in Medicinal Chemistry from Jagiellonian University. Before moving fully into software engineering, I worked across drug-discovery data science, RDKit-based ligand generation, bioinorganic chemistry, and photodynamic-therapy research, including an academic project at the University of Glasgow.
This scientific background shapes how I approach software: chemical assumptions, data provenance, and experimental context should remain visible in the implementation.
About this blog
I use cheminfo.dev to document what I learn while building scientific software.
The posts focus on practical cheminformatics: molecular representations, chemical databases, data-quality problems, RDKit workflows, open-source tools, and the engineering decisions behind reproducible computational chemistry.
The aim is to publish useful technical notes—not polished marketing material—and to explain both what worked and what did not.