Skip to content

Fix FP8 current scaling FakeTensor output shape - #3659

Open
sanjana658 wants to merge 2 commits into
NVIDIA:mainfrom
sanjana658:fix/3636-fp8-fake-scale-shape
Open

sanjana658 wants to merge 2 commits into
NVIDIA:mainfrom
sanjana658:fix/3636-fp8-fake-scale-shape

Conversation

@sanjana658

Copy link
Copy Markdown

Description

Please include a brief summary of the changes, relevant motivation and context.

Fixes # (issue)

Summary

  • Return a scalar inverse-scale tensor from the fake implementation of tex::fp8_cs_quantize, matching the eager output metadata.
  • Add a regression test using torch.library.opcheck with test_faketensor.

Validation

  • Python syntax compilation passed for both changed files.
  • git diff --check passed.
  • Native CUDA opcheck test has not been run locally because the local environment has no CUDA GPU and TransformerEngine's compiled package is not installed.

Related issue: #3636

@sanjana658
sanjana658 requested a review from ksivaman as a code owner October 9, 2026 02:48
@github-actions github-actions Bot added the community-contribution PRs from external contributor outside the core maintainers, representing community-driven work. label Oct 9, 2026
@greptile-apps

greptile-apps Bot commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[Medium impact] The PR appears safe to merge; the fake inverse scale now matches the eager output.

Summary

Makes the fake output of tex::fp8_cs_quantize return a scalar inverse scale, matching eager execution.

  • Adds a focused torch.library.opcheck regression test.
  • Adds the test to the L0 PyTorch runner, addressing the earlier CI coverage comment.
  • No new actionable issues were found. CUDA execution was not verified during this review.

Reviews (3) · Last reviewed commit: "test: include FP8 fake tensor opcheck in..." · Reviewed by Greptile

Comment thread tests/pytorch/test_fp8_cs_quantize_opcheck.py
Sanjana Thanabalan added 2 commits October 9, 2026 08:28
Signed-off-by: Sanjana Thanabalan <190764683+sanjana658@users.noreply.github.com>
Signed-off-by: Sanjana Thanabalan <190764683+sanjana658@users.noreply.github.com>
@sanjana658
sanjana658 force-pushed the fix/3636-fp8-fake-scale-shape branch from 604ab7b to 019b153 Compare October 9, 2026 03:00
@sanjana658

Copy link
Copy Markdown
Author

Hi maintainers, I've addressed the CI coverage feedback by adding the regression test to the L0 PyTorch test runner, and the DCO check is now passing. The PR is ready for review. Could someone please approve the pending workflows and review the changes when convenient? Thank you!

@@ -0,0 +1,17 @@
# Copyright (c) 2022-2026, NVIDIA CORPORATION & AFFILIATES. All rights reserved.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please add this test to the existing test_onnx_export.py instead rather than creating a completely new file.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants