Skip to content

Raisetolinalg - #412

Draft
arpitj1 wants to merge 184 commits into
llvm:raisetolinalgfrom
arpitj1:raisetolinalg
Draft

Raisetolinalg#412
arpitj1 wants to merge 184 commits into
llvm:raisetolinalgfrom
arpitj1:raisetolinalg

Conversation

@arpitj1

@arpitj1 arpitj1 commented Jun 6, 2024

Copy link
Copy Markdown
Collaborator

Some modifications to fuse linalg.generic op with for op

arpitj1 added 30 commits June 6, 2024 08:41
…f debufferizing added which works for tiling and fusion
arpitj1 added 30 commits June 1, 2026 18:19
Lower non-injective submap writebacks correctly, outline residual stages for GPU execution, add persistent workspace and CUDA graph support, compose cuTensorNet networks, and restrict matcher emission to ABI-lowerable implementations. Include focused regression coverage and updated correctness tooling.
Add shape-matched native and raised silicon measurements, residency analysis inputs, generation scripts, and the corrected adaptive-average-pool match artifact used by the numerical viewer.
Preserve exact H(curl)/H(div) scratch extents, split unsafe multi-output residual reductions, correct native timing synchronization, and record regenerated matcher, application, CUDA graph, correctness, and Orin performance evidence.
RMSNorm no longer has a runtime shim, so stop treating it as a positive lowering case and add a regression requiring unsupported semantic-only symbols to fail explicitly.
Verify that C + B * (A * alpha) is recognized as cublasDgemm and lowered to the runtime ABI, while documenting that this exact expression requires commutative reordering rather than associative Egglog saturation.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants