Repository navigation
Conversation
|
Thanks for your PR! I think you also need to update the |
|
Yeah I'm currently fixing it. |
There was a problem hiding this comment.
Pull request overview
Adds a FlashAttention example kernel implemented in Allo and wires it into CI so the example is exercised during PR runs.
Changes:
- Added a FlashAttention Allo schedule/kernel implementation (
get_scheduled_flash_attention). - Added a Python testbench that builds/runs the kernel and compares against a NumPy golden reference.
- Updated GitHub Actions workflow to execute the FlashAttention testbench.
Reviewed changes
Copilot reviewed 3 out of 3 changed files in this pull request and generated 6 comments.
| File | Description |
|---|---|
examples/flashattention/test_flash.py |
Testbench that computes a golden attention output and validates the Allo-generated kernel (plus optional Vitis HLS synthesis). |
examples/flashattention/flash_Atten.py |
Allo implementation/schedule for a tiled FlashAttention-style attention computation. |
.github/workflows/config.yml |
Adds a CI step to run the FlashAttention testbench. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
You can also share your feedback on Copilot code review. Take the survey.
chhzh123
left a comment
There was a problem hiding this comment.
Can you fix Copilot's suggestions (if they are valid)? Also, it'd be good to attach your performance results in the PR description
Fangtangtang
left a comment
There was a problem hiding this comment.
Hi @RuizeYu05, sorry for the very delayed review.
Shall we put falshattention in this PR and MHA in #581 under the same directory named "attention"? It would also be great to add a README under this directory describing the kernels (e.g., algorithm/dataflow/performance results).
Also, the current CI only runs the simulator. Could you add this example to .github/workflows/fpga_weekly.yml to enable tests on hls?

Description
I add the example of implementing flash attention using allo and modify the CI test to support the testing of my flash attention kernel.
Problems
Implementing flash attention using allo
Proposed Solutions
Add an example to the application of allo
Examples
Checklist
Please make sure to review and check all of these items: