Optimize triangles - #55
Conversation
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #55 +/- ##
==========================================
+ Coverage 99.37% 99.45% +0.08%
==========================================
Files 5 5
Lines 161 185 +24
==========================================
+ Hits 160 184 +24
Misses 1 1
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Sentry. 🚀 New features to boost your workflow:
|
will require downstream changes in expected
e289665 to
7f8ece8
Compare
|
Ready for review Just a note, the To make triangles for one image,
This logic will be easier to implement with the explicit loop architecture I use in |
|
This is brilliant, thanks Chris! I am also going to use this as an opportunity for me to finally learn about AirspeedVeclocity.jl and implement it in this repo. Once ConsensusFitting.jl is registered, I think the PR here would be a great first test case, so I will just leave it open until then. btw, I am seeing similar speed-ups on my end too 👍🏾 ┌────────────────────────┬─────────────┬────────────┬─────────┐
│ Benchmark │ Median Time │ Memory │ Allocs │
├────────────────────────┼─────────────┼────────────┼─────────┤
│ _build_correspondences │ 159.065 ms │ 143.37 MiB │ 2910662 │
│ _triangle_invariants │ 15.031 ms │ 39.48 MiB │ 808508 │
└────────────────────────┴─────────────┴────────────┴─────────┘
┌────────────────────────┬─────────────┬───────────┬────────┐
│ Benchmark │ Median Time │ Memory │ Allocs │
├────────────────────────┼─────────────┼───────────┼────────┤
│ _build_correspondences │ 87.688 ms │ 34.80 MiB │ 323439 │
│ _triangle_invariants │ 1.651 ms │ 6.17 MiB │ 6 │
└────────────────────────┴─────────────┴───────────┴────────┘ |
|
Sounds good, FYI I think you will need to add an |
|
You could even get into stuff like @inline function _canonical_vertex_order(i, j, l, d2_ji, d2_lj, d2_li, xs, ys)
# There are three possible longest edges:
# 1. ji is longest (c1 == true)
# 2. lj is longest (c2 == true)
# 3. li is longest (neither c1 nor c2)
ji_longest = (d2_ji >= d2_lj) & (d2_ji >= d2_li)
lj_longest = (d2_lj >= d2_ji) & (d2_lj >= d2_li)
# The nested ifelse expressions are equivalent to:
#
# if c1
# v1, v2, apex = i, j, l
# elseif c2
# v1, v2, apex = j, l, i
# else
# v1, v2, apex = i, l, j
# end
# Using ifelse keeps this branchless
v1 = ifelse(ji_longest, i, ifelse(lj_longest, j, i))
v2 = ifelse(ji_longest, j, ifelse(lj_longest, l, l))
apex = ifelse(ji_longest, l, ifelse(lj_longest, i, j))
# Enforce CCW winding: swap base vertices if cross product is negative
@inbounds cross = (xs[v2] - xs[v1]) * (ys[apex] - ys[v1]) - (ys[v2] - ys[v1]) * (xs[apex] - xs[v1])
swap = cross < 0
vv1 = ifelse(swap, v2, v1)
vv2 = ifelse(swap, v1, v2)
return vv1, vv2, apex
endThis emits better LLVM and results in speedup for |
|
Thanks Chris, just left my very light review responses above.
Thanks for the heads up! I'm also getting up to speed on what gaps might still exist with
That's a good concern. I think just having good docs and/or docstrings for things like this would probably be enough if you think it would be worth it for the performance improvement |
If you copy-paste the revised function definitions from this PR into
your REPL you can run the following comparison which will show the new
implementation is ~3x faster and allocation free **EDIT: If you have to
collect the input (as we do because of how Photometry.jl passes us the
cutout) then it will be 2 allocations and adds ~100 ns to the base case
speed**. I chose to do `com_psf(T::Type{<:AbstractFloat}, img_ap,
rel_thresh)` because Float64 is not really any slower on my testing so
this way it makes it easier for us to switch to Float64 in the future if
we every wanted to.
```julia
using PSFModels: gaussian
using BenchmarkTools
import Astroalign
const x = 1:20
const y = 1:20
const T = Float32
model(x, y, amp) = gaussian(T, 4, 4; x, y, fwhm=3) * amp
data = model.(x, y', 10) .+ T(0.1) * randn(T, length(x), length(y))
result_orig = Astroalign.com_psf(data; rel_thresh=0.1f0)
result_revised = com_psf(data; rel_thresh=0.1f0)
println("Original COM: ", (result_orig.psf_params.x, result_orig.psf_params.y))
println("Revised COM: ", (result_revised.psf_params.x, result_revised.psf_params.y))
println("Original FWHM: ", result_orig.psf_params.fwhm)
println("Revised FWHM: ", result_revised.psf_params.fwhm)
println("Original benchmark: ")
display(@benchmark Astroalign.com_psf($data; rel_thresh=$0.1f0))
println("Revised benchmark: ")
display(@benchmark com_psf($data; rel_thresh=$0.1f0))
```
My result:
**EDIT:** Because Photometry.jl will pass us an object from
Transducers.jl we have to call `collect`, giving us 2 allocations and
runtime +100 ns over not collecting (if we had a pure matrix input).
```julia
julia> println("Original COM: ", (result_orig.psf_params.x, result_orig.psf_params.y))
Original COM: (3.9974918f0, 3.9858212f0)
julia> println("Revised COM: ", (result_revised.psf_params.x, result_revised.psf_params.y))
Revised COM: (3.9974945f0, 3.9858336f0)
julia> println("Original FWHM: ", result_orig.psf_params.fwhm)
Original FWHM: (2.31581660320762, 2.3766457470426725)
julia> println("Revised FWHM: ", result_revised.psf_params.fwhm)
Revised FWHM: (2.3164616f0, 2.3772223f0)
julia> display(@benchmark Astroalign.com_psf($data; rel_thresh=$0.1f0))
BenchmarkTools.Trial: 10000 samples with 10 evaluations per sample.
Range (min … max): 1.408 μs … 1.050 ms ┊ GC (min … max): 0.00% … 99.19%
Time (median): 1.544 μs ┊ GC (median): 0.00%
Time (mean ± σ): 2.372 μs ± 15.628 μs ┊ GC (mean ± σ): 21.39% ± 4.62%
▆█▇▅▃▁▁▁ ▃▄▃▂ ▂
█████████▇▇▆▇▇▇▇▆▇▆▆▆▇▆██████▆▆▆▅▄▅▄▅▅▃▅▆███▇▅▄▄▄▁▁▄▁▆▅▇██ █
1.41 μs Histogram: log(frequency) by time 5.56 μs <
Memory estimate: 10.41 KiB, allocs estimate: 20.
julia> println("Revised benchmark: ")
Revised benchmark:
julia> @benchmark com_psf($data; rel_thresh=$0.1f0)
BenchmarkTools.Trial: 10000 samples with 183 evaluations per sample.
Range (min … max): 571.716 ns … 9.490 μs ┊ GC (min … max): 0.00% … 88.05%
Time (median): 600.563 ns ┊ GC (median): 0.00%
Time (mean ± σ): 678.928 ns ± 532.693 ns ┊ GC (mean ± σ): 8.48% ± 9.61%
█▅▂▁▁ ▁
██████▇▇▆▆▄▅▁▄▃▁▃▁▁▁▁▁▃▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▃▁▃▄▃▄▄▄▄▄▄▄▃▄▅ █
572 ns Histogram: log(frequency) by time 4.5 μs <
Memory estimate: 1.64 KiB, allocs estimate: 2.
```
---------
Co-authored-by: Ian Weaver <weaveric@gmail.com>
Optimizes the building of the correspondences and triangles to minimize the impact of that step in the case that it is repeated in workflows when fitting multiple images (#52). Also added a benchmark script for you to compare if you'd like. Note that comparisons to main are not totally fair as
canonical_vertex_orderis now called internally to_triangle_invariantswhereas before it was an additional post-processing step (so there's actually more work being done in_triangle_invariantsthan before. Even so,_triangle_invariantsfor 100 stars now takes ~1.5 ms on my machine compared to 10.9 ms onmain.Combinatorics.combinationsiterator with explicit triple-nestedforloops over strictly increasing index triples(i < j < l), eliminating iterator allocation.Distances.euclideanwith inlined squared-distance calculation, avoidingsqrtuntil the final ratio computation (the sqrt may not even be strictly necessary, we could do matching in distance^2 space, but I digress)sort!([...])per triangle with a branchless 3-element sorting networkmap+stackresult construction with a single pre-allocatedMatrix{Float64}written in-place, eliminating intermediate allocations and a second pass over the data_canonical_vertex_orderdirectly into_triangle_invariants, soCis returned in canonical vertex order (apex last, base vertices CCW) without a separate post-processing pass; the squared distances already computed in the loop are reused directlyCombinatorics,Distancesdependencies