Skip to content

Commit 322ef2b

Browse files
committed
0.0.7
1 parent cfc1d6d commit 322ef2b

19 files changed

Lines changed: 3090 additions & 60 deletions

‎README.md‎

Lines changed: 33 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -1,9 +1,10 @@
1-
# PyImageCUDA 0.0.6
1+
# PyImageCUDA 0.0.7
22

33
[![PyPI version](https://img.shields.io/pypi/v/pyimagecuda.svg)](https://pypi.org/project/pyimagecuda/)
44
[![Build Status](https://github.com/offerrall/pyimagecuda/actions/workflows/build.yml/badge.svg)](https://github.com/offerrall/pyimagecuda/actions)
55
![Python](https://img.shields.io/badge/python-3.10%20|%203.11%20|%203.12%20|%203.13-blue)
66
![Platform](https://img.shields.io/badge/platform-Windows%20%7C%20Linux-brightgreen)
7+
![Tests](https://img.shields.io/badge/tests-85%20passed-brightgreen)
78
[![NVIDIA](https://img.shields.io/badge/NVIDIA-CUDA-76B900?style=flat&logo=nvidia&logoColor=white)](https://developer.nvidia.com/cuda-zone)
89

910
**GPU-accelerated image compositing for Python.**
@@ -74,6 +75,17 @@ pip install pyimagecuda
7475
* [Filter](https://offerrall.github.io/pyimagecuda/filter/) (Gaussian Blur, Sharpen, Sepia, Invert, Threshold, Solarize, Sobel, Emboss)
7576
* [Effect](https://offerrall.github.io/pyimagecuda/effect/) (Drop Shadow, Rounded Corners, Stroke, Vignette)
7677

78+
## Performance
79+
80+
PyImageCUDA shows significant speedups for GPU-friendly operations like blending, filtering, and transformations. Performance varies by operation complexity and workflow:
81+
82+
- Complex operations (blur, blend, rotate) see **10-260x improvements**
83+
- Simple operations (flip, crop) see **3-20x improvements**
84+
- Real-world pipelines with file I/O typically see **1.5-2.5x speedups**
85+
86+
Results depend on your hardware, batch size, and whether you reuse GPU buffers.
87+
88+
**[→ View Detailed Benchmarks](https://offerrall.github.io/pyimagecuda/benchmarks/)**
7789

7890
## Requirements
7991

@@ -85,8 +97,27 @@ pip install pyimagecuda
8597

8698
**NOT REQUIRED:** Visual Studio, CUDA Toolkit, or Conda.
8799

100+
## Linux Compatibility & Troubleshooting
101+
102+
PyImageCUDA is currently tested primarily on **Ubuntu LTS** releases with up-to-date NVIDIA drivers.
103+
104+
If you encounter the following error on Linux:
105+
106+
```text
107+
RuntimeError: Kernel launch failed: the provided PTX was compiled with an unsupported toolchain.
108+
```
109+
110+
Solution: This indicates your installed NVIDIA drivers are too old to execute the kernels included in the library. Please update your NVIDIA drivers to the latest version available for your distribution (Proprietary drivers recommended).
111+
112+
We are actively investigating ways to broaden compatibility for older drivers and legacy Linux distributions in future releases.
113+
114+
## Tests
115+
```bash
116+
pytest tests/tests.py
117+
```
118+
88119
## Contributing
89120
Contributions welcome! Open issues or submit PRs
90121

91122
## License
92-
MIT License. See [LICENSE](LICENSE) for details.
123+
MIT License. See [LICENSE](LICENSE) for details.

‎benchmarks/BENCHMARK_REPORT.md‎

Lines changed: 149 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,149 @@
1+
# PyImageCUDA Performance Report
2+
Generated automatically by /benchmarks/benchmarks.py
3+
4+
## Gaussian Blur Benchmark (1080p)
5+
> **Config:** Image: `photo.jpg`, Radius: `20`, Iterations: `50`
6+
7+
### Pure Algorithm (Compute Bound)
8+
| Library | Avg (ms) | FPS | Speedup |
9+
| :--- | :--- | :--- | :--- |
10+
| **PyImageCUDA (Reuse)** | 0.97 | 1032.7 | 54.4x |
11+
| **PyImageCUDA (Alloc)** | 3.49 | 286.4 | 15.1x |
12+
| **OpenCV** | 50.85 | 19.7 | 1.0x |
13+
| **Pillow** | 52.67 | 19.0 | 1.0x |
14+
15+
### End-to-End (Disk I/O + Encode)
16+
| Library | Avg (ms) | FPS | Speedup |
17+
| :--- | :--- | :--- | :--- |
18+
| **PyImageCUDA E2E (Buffered)** | 83.72 | 11.9 | 2.7x |
19+
| **PyImageCUDA E2E** | 92.06 | 10.9 | 2.4x |
20+
| **OpenCV E2E** | 100.19 | 10.0 | 2.2x |
21+
| **Pillow E2E** | 224.84 | 4.4 | 1.0x |
22+
23+
---
24+
25+
## Blend Normal Benchmark (1080p)
26+
> **Config:** Image: `photo.jpg`, Iterations: `50`
27+
28+
### Pure Algorithm (Compute Bound)
29+
| Library | Avg (ms) | FPS | Speedup |
30+
| :--- | :--- | :--- | :--- |
31+
| **PyImageCUDA** | 0.27 | 3705.1 | 376.7x |
32+
| **Pillow** | 12.55 | 79.7 | 8.1x |
33+
| **OpenCV** | 101.67 | 9.8 | 1.0x |
34+
35+
### End-to-End (Disk I/O + Encode)
36+
| Library | Avg (ms) | FPS | Speedup |
37+
| :--- | :--- | :--- | :--- |
38+
| **PyImageCUDA E2E (Buffered)** | 153.80 | 6.5 | 2.2x |
39+
| **PyImageCUDA E2E** | 159.17 | 6.3 | 2.1x |
40+
| **OpenCV E2E** | 242.98 | 4.1 | 1.4x |
41+
| **Pillow E2E** | 334.47 | 3.0 | 1.0x |
42+
43+
---
44+
45+
## Resize Bilinear Benchmark (1080p -> 800x600)
46+
> **Config:** Image: `photo.jpg`, Target: `800x600`, Interpolation: `Bilinear`, Iterations: `50`
47+
48+
### Pure Algorithm (Compute Bound)
49+
| Library | Avg (ms) | FPS | Speedup |
50+
| :--- | :--- | :--- | :--- |
51+
| **PyImageCUDA Bilinear (Reuse)** | 0.14 | 7096.3 | 132.9x |
52+
| **OpenCV Bilinear** | 0.52 | 1940.0 | 36.3x |
53+
| **PyImageCUDA Bilinear (Alloc)** | 0.80 | 1249.7 | 23.4x |
54+
| **Pillow Bilinear** | 18.73 | 53.4 | 1.0x |
55+
56+
### End-to-End (Disk I/O + Encode)
57+
| Library | Avg (ms) | FPS | Speedup |
58+
| :--- | :--- | :--- | :--- |
59+
| **OpenCV Bilinear E2E** | 40.42 | 24.7 | 2.5x |
60+
| **PyImageCUDA Bilinear E2E (Buffered)** | 49.53 | 20.2 | 2.0x |
61+
| **PyImageCUDA Bilinear E2E** | 52.37 | 19.1 | 1.9x |
62+
| **Pillow Bilinear E2E** | 99.91 | 10.0 | 1.0x |
63+
64+
---
65+
66+
## Resize Lanczos Benchmark (1080p -> 800x600)
67+
> **Config:** Image: `photo.jpg`, Target: `800x600`, Interpolation: `Lanczos/Bicubic`, Iterations: `50`
68+
69+
### Pure Algorithm (Compute Bound)
70+
| Library | Avg (ms) | FPS | Speedup |
71+
| :--- | :--- | :--- | :--- |
72+
| **PyImageCUDA Lanczos (Reuse)** | 0.88 | 1131.3 | 35.2x |
73+
| **PyImageCUDA Lanczos (Alloc)** | 1.62 | 617.8 | 19.2x |
74+
| **OpenCV Lanczos** | 4.16 | 240.2 | 7.5x |
75+
| **Pillow Lanczos** | 31.15 | 32.1 | 1.0x |
76+
77+
### End-to-End (Disk I/O + Encode)
78+
| Library | Avg (ms) | FPS | Speedup |
79+
| :--- | :--- | :--- | :--- |
80+
| **OpenCV Lanczos E2E** | 44.42 | 22.5 | 2.5x |
81+
| **PyImageCUDA Lanczos E2E (Buffered)** | 52.28 | 19.1 | 2.1x |
82+
| **PyImageCUDA Lanczos E2E** | 54.43 | 18.4 | 2.0x |
83+
| **Pillow Lanczos E2E** | 108.93 | 9.2 | 1.0x |
84+
85+
---
86+
87+
## Rotate 35° Benchmark (1080p)
88+
> **Config:** Image: `photo.jpg`, Angle: `35°`, Expand: `True`, Iterations: `50`
89+
90+
### Pure Algorithm (Compute Bound)
91+
| Library | Avg (ms) | FPS | Speedup |
92+
| :--- | :--- | :--- | :--- |
93+
| **PyImageCUDA (Reuse)** | 0.30 | 3336.0 | 260.6x |
94+
| **PyImageCUDA (Alloc)** | 3.31 | 301.9 | 23.6x |
95+
| **OpenCV** | 5.98 | 167.2 | 13.1x |
96+
| **Pillow** | 78.11 | 12.8 | 1.0x |
97+
98+
### End-to-End (Disk I/O + Encode)
99+
| Library | Avg (ms) | FPS | Speedup |
100+
| :--- | :--- | :--- | :--- |
101+
| **OpenCV E2E** | 156.55 | 6.4 | 2.8x |
102+
| **PyImageCUDA E2E** | 162.23 | 6.2 | 2.7x |
103+
| **PyImageCUDA E2E (Buffered)** | 162.33 | 6.2 | 2.7x |
104+
| **Pillow E2E** | 432.76 | 2.3 | 1.0x |
105+
106+
---
107+
108+
## Flip Horizontal Benchmark (1080p)
109+
> **Config:** Image: `photo.jpg`, Direction: `Horizontal`, Iterations: `50`
110+
111+
### Pure Algorithm (Compute Bound)
112+
| Library | Avg (ms) | FPS | Speedup |
113+
| :--- | :--- | :--- | :--- |
114+
| **PyImageCUDA (Reuse)** | 0.17 | 5916.0 | 20.3x |
115+
| **PyImageCUDA (Alloc)** | 1.73 | 577.5 | 2.0x |
116+
| **OpenCV** | 3.12 | 320.6 | 1.1x |
117+
| **Pillow** | 3.43 | 291.7 | 1.0x |
118+
119+
### End-to-End (Disk I/O + Encode)
120+
| Library | Avg (ms) | FPS | Speedup |
121+
| :--- | :--- | :--- | :--- |
122+
| **OpenCV E2E** | 126.09 | 7.9 | 2.6x |
123+
| **PyImageCUDA E2E (Buffered)** | 149.56 | 6.7 | 2.2x |
124+
| **PyImageCUDA E2E** | 150.95 | 6.6 | 2.2x |
125+
| **Pillow E2E** | 324.85 | 3.1 | 1.0x |
126+
127+
---
128+
129+
## Crop Center Benchmark (1080p → 512×512)
130+
> **Config:** Image: `photo.jpg`, Source: `1920×1080`, Output: `512×512`, Iterations: `50`
131+
132+
### Pure Algorithm (Compute Bound)
133+
| Library | Avg (ms) | FPS | Speedup |
134+
| :--- | :--- | :--- | :--- |
135+
| **PyImageCUDA (Reuse)** | 0.04 | 27037.3 | 13.3x |
136+
| **OpenCV** | 0.25 | 3987.1 | 2.0x |
137+
| **Pillow** | 0.27 | 3692.7 | 1.8x |
138+
| **PyImageCUDA (Alloc)** | 0.49 | 2029.6 | 1.0x |
139+
140+
### End-to-End (Disk I/O + Encode)
141+
| Library | Avg (ms) | FPS | Speedup |
142+
| :--- | :--- | :--- | :--- |
143+
| **OpenCV E2E** | 27.37 | 36.5 | 2.0x |
144+
| **PyImageCUDA E2E (Buffered)** | 35.72 | 28.0 | 1.6x |
145+
| **PyImageCUDA E2E** | 38.34 | 26.1 | 1.5x |
146+
| **Pillow E2E** | 55.69 | 18.0 | 1.0x |
147+
148+
---
149+

0 commit comments

Comments
 (0)