Skip to content

Synchronize before indexing host-accessible arrays from the CPU - #649

Merged
maleadt merged 1 commit into
mainfrom
tb/host-indexing-sync
Sep 30, 2026
Merged

maleadt merged 1 commit into
mainfrom
tb/host-indexing-sync

Conversation

@maleadt

@maleadt maleadt commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

Arrays backed by shared or host memory can be indexed directly from the CPU, but doing so didn't wait for work still queued on the GPU. The most visible consequence is that reductions of such arrays returned garbage, because the result was read before the kernel computing it had finished:

a = oneArray{Float32, 1, oneL0.SharedBuffer}(fill(1f0, 1024))
sum(a)       # returned 0 instead of 1024
maximum(a)   # returned 0 instead of 1

Scalar indexing of these arrays now synchronizes first, so this returns the right results. As with other implicit synchronization, this only waits for the current task's queue: work submitted from other tasks, or to a queue you created yourself, still needs to be synchronized explicitly.

Found while working on host-memory wrapping for oneAPI.jl, but independent of it. Tested on an Iris Xe: the new test fails on main (sum and maximum return 0 for both shared and host buffers) and passes with this change.

Scalar getindex and setindex! on shared and host buffers accessed the memory directly,
without waiting for queued work. Reading the result of a reduction therefore returned
stale data, e.g. sum() and maximum() of such arrays returned 0. Synchronize the current
task's stream first.
@github-actions

Copy link
Copy Markdown
Contributor

Your PR requires formatting changes to meet the project's style guidelines.
Please consider running Runic (git runic main) to apply these changes.

Click here to view the suggested changes.
diff --git a/test/array.jl b/test/array.jl
index eb21d26..e386d9e 100644
--- a/test/array.jl
+++ b/test/array.jl
@@ -104,11 +104,11 @@ end
 end
 
 @testset "reductions of host-accessible arrays" begin
-  for B in (oneL0.SharedBuffer, oneL0.HostBuffer)
-    a = oneArray{Float32, 1, B}(fill(1.0f0, 1024))
-    @test sum(a) == 1024
-    @test maximum(a) == 1
-  end
+    for B in (oneL0.SharedBuffer, oneL0.HostBuffer)
+        a = oneArray{Float32, 1, B}(fill(1.0f0, 1024))
+        @test sum(a) == 1024
+        @test maximum(a) == 1
+    end
 end
 
 # https://github.com/JuliaGPU/CUDA.jl/issues/2191

@codecov

codecov Bot commented Sep 29, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 78.96%. Comparing base (85dd4cf) to head (5c33db5).

Additional details and impacted files
@@            Coverage Diff             @@
##             main     #649      +/-   ##
==========================================
- Coverage   80.87%   78.96%   -1.91%     
==========================================
  Files          56       56              
  Lines        4088     4089       +1     
==========================================
- Hits         3306     3229      -77     
- Misses        782      860      +78     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@maleadt
maleadt merged commit 5056aa9 into main Sep 30, 2026
4 of 5 checks passed
@maleadt
maleadt deleted the tb/host-indexing-sync branch September 30, 2026 04:50
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant