Metadata-Version: 2.5
Name: faster-diffbloch
Version: 0.1.2
Summary: Drop-in Metal GPU (macOS) and optimized CPU (macOS/Linux) acceleration for diffBloch
Project-URL: Homepage, https://godofecht.github.io/diffFlow/
Project-URL: Documentation, https://godofecht.github.io/diffFlow/
Project-URL: Repository, https://github.com/godofecht/diffFlow
Project-URL: Issues, https://github.com/godofecht/diffFlow/issues
Project-URL: PyPI, https://pypi.org/project/faster-diffbloch/
Project-URL: piwheels, https://www.piwheels.org/project/faster-diffbloch/
Project-URL: Original diffBloch, https://diffbloch.com
Author-email: Abhishek Shivakumar <abhishek@example.com>
License-Expression: MIT
License-File: LICENSE
Keywords: Metal GPU,PyTorch acceleration,diffBloch,electron crystallography,electron diffraction,structure refinement
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: MacOS
Classifier: Operating System :: MacOS :: MacOS X
Classifier: Operating System :: POSIX :: Linux
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Physics
Requires-Python: >=3.10
Requires-Dist: numpy>=1.24
Requires-Dist: torch>=2.0
Provides-Extra: diffbloch
Requires-Dist: diffbloch; extra == 'diffbloch'
Description-Content-Type: text/markdown

# faster-diffbloch

Drop-in Apple Silicon Metal GPU and optimized CPU acceleration for [diffBloch](https://diffbloch.com) electron crystallography structure refinement.

**Up to 1.80× faster than PyTorch CPU** in measured end-to-end forward-plus-backward benchmarks, with a Metal GPU path for Apple Silicon and an optimized CPU path for macOS and Linux.

| Package | Documentation | Source repository | Original project |
| :--- | :--- | :--- | :--- |
| [PyPI](https://pypi.org/project/faster-diffbloch/) · [piwheels](https://www.piwheels.org/project/faster-diffbloch/) | [diffFlow documentation and benchmarks](https://godofecht.github.io/diffFlow/) | [github.com/godofecht/diffFlow](https://github.com/godofecht/diffFlow) | [diffbloch.com](https://diffbloch.com) |

---

## Operating System and Platform Support

| Operating System / Hardware | CPU Acceleration (`device="cpu"`) | GPU Acceleration (`device="gpu"`) | Backend Runtime |
| :--- | :---: | :---: | :--- |
| **macOS Apple Silicon (M1/M2/M3/M4/Max/Ultra)** | **Supported** | **Supported** | Native Metal Compute Shaders + Apple Accelerate BLAS |
| **macOS Intel (x86_64)** | **Supported** | Fallback to CPU | Apple Accelerate BLAS |
| **Linux (x86_64 / aarch64)** | **Supported** | Fallback to CPU | OpenBLAS / C11 BLAS |
| **Windows** | PyTorch Reference | PyTorch Reference | Pure PyTorch Reference Fallback |

Runtime platform guards automatically detect your operating system and hardware configuration. When `device="gpu"` is requested on Linux, `faster-diffbloch` automatically selects the optimized CPU backend with an informative warning.

---

## Why faster-diffBloch?

`faster-diffbloch` keeps the diffBloch workflow and public API intact while moving
the expensive propagation and gradient work onto an accelerated native backend.
You get the same scientific calculation, with a faster path on supported hardware
and a safe PyTorch fallback when the native backend is unavailable.

The package is validated against the diffBloch test suite and reference results.
It passes all 738 diffBloch tests, 55 additional conformance tests, and reproduces
the experimental quartz refinement result.

---

## Performance

End-to-end forward-plus-backward timing on an Apple M4 Max, measured across the
same diffBloch workload at different beam counts:

| Beams | faster-diffBloch | PyTorch | Speedup |
| :---: | :---: | :---: | :---: |
| 31 | 0.90 ms | 1.62 ms | **1.80×** |
| 61 | 1.37 ms | 2.24 ms | **1.63×** |
| 91 | 2.84 ms | 3.99 ms | **1.41×** |
| 163 | 7.67 ms | 10.98 ms | **1.43×** |
| 579 | 89.81 ms | 135.77 ms | **1.51×** |

That is **1.41×–1.80× faster than PyTorch CPU** across the measured cases,
including the largest 579-beam case. The same benchmark is **1.21× faster than
the Mojo reference at 31 beams**, remains ahead through 163 beams, and is
essentially level at 579 beams.

On Apple Silicon, the Metal GPU path is also available. In the package's M4 Max
comparison run at 579 beams, it measured **2.24× faster than PyTorch CPU for the
forward pass** and **2.63× faster than PyTorch MPS for forward plus backward**.

These figures are workload benchmarks, not a promise that every crystal,
hardware configuration, or beam count will see the same result. The CPU table
uses single-threaded PyTorch and the same Apple Accelerate environment for a
like-for-like comparison.

### Package benchmark snapshot

| Implementation | Forward | Forward + Backward | Speedup vs PyTorch CPU | Speedup vs PyTorch MPS |
| :--- | :---: | :---: | :---: | :---: |
| PyTorch CPU | 25.7 ms | 130.3 ms | 1.00x | 1.17x |
| PyTorch MPS (fallback) | 26.2 ms | 153.0 ms | 0.85x | 1.00x |
| **faster-diffBloch CPU** | **24.4 ms** | **83.5 ms** | **1.56x** | **1.83x** |
| **faster-diffBloch Metal GPU** | **13.1 ms** | **58.1 ms** | **2.24x** | **2.63x** |

---

## Installation

```bash
pip install faster-diffbloch
```

---

## Usage

### 1. Drop-in CLI

Use `diffbloch-fast` or `faster-diffbloch` anywhere you would use `diffbloch`:

```bash
diffbloch-fast infer examples/Colmey_et_al_2026/data/quartz-no-abs
diffbloch-fast refine examples/Colmey_et_al_2026/data/quartz-no-abs
```

### 2. Python API Injection

Enable acceleration inside any existing diffBloch script:

```python
import faster_diffbloch

# Enable Metal GPU acceleration (macOS Apple Silicon)
faster_diffbloch.enable(device="gpu")

# Or CPU acceleration (macOS and Linux)
faster_diffbloch.enable(device="cpu")

# Run standard diffBloch code
import diffBloch
# All propagate and matrix_exp calls now route through faster-diffBloch
```
