Logo
  • 1. Introduction
    • 1.1. Scalable Data-Parallel Computing using GPUs
    • 1.2. Goals of PTX
    • 1.3. PTX ISA Version 9.4
    • 1.4. Document Structure
  • 2. Programming Model
    • 2.1. A Highly Multithreaded Coprocessor
    • 2.2. Thread Hierarchy
      • 2.2.1. Cooperative Thread Arrays
      • 2.2.2. Cluster of Cooperative Thread Arrays
      • 2.2.3. Grid of Clusters
    • 2.3. Memory Hierarchy
  • 3. PTX Machine Model
    • 3.1. A Set of SIMT Multiprocessors
    • 3.2. Independent Thread Scheduling
    • 3.3. On-chip Shared Memory
  • 4. Syntax
    • 4.1. Source Format
    • 4.2. Comments
    • 4.3. Statements
      • 4.3.1. Directive Statements
      • 4.3.2. Instruction Statements
    • 4.4. Identifiers
    • 4.5. Constants
      • 4.5.1. Integer Constants
      • 4.5.2. Floating-Point Constants
      • 4.5.3. Predicate Constants
      • 4.5.4. Constant Expressions
      • 4.5.5. Integer Constant Expression Evaluation
      • 4.5.6. Summary of Constant Expression Evaluation Rules
  • 5. State Spaces, Types, and Variables
    • 5.1. State Spaces
      • 5.1.1. Register State Space
      • 5.1.2. Special Register State Space
      • 5.1.3. Constant State Space
        • 5.1.3.1. Banked Constant State Space (deprecated)
      • 5.1.4. Global State Space
      • 5.1.5. Local State Space
      • 5.1.6. Parameter State Space
        • 5.1.6.1. Kernel Function Parameters
        • 5.1.6.2. Kernel Function Parameter Attributes
        • 5.1.6.3. Kernel Parameter Attribute: .ptr
        • 5.1.6.4. Device Function Parameters
      • 5.1.7. Shared State Space
      • 5.1.8. Texture State Space (deprecated)
    • 5.2. Types
      • 5.2.1. Fundamental Types
      • 5.2.2. Restricted Use of Sub-Word Sizes
      • 5.2.3. Alternate Floating-Point Data Formats
      • 5.2.4. Fixed-point Data format
      • 5.2.5. Packed Data Types
        • 5.2.5.1. Packed Floating Point Data Types
        • 5.2.5.2. Packed Integer Data Types
        • 5.2.5.3. Packed Fixed-Point Data Types
    • 5.3. Texture Sampler and Surface Types
      • 5.3.1. Texture and Surface Properties
      • 5.3.2. Sampler Properties
      • 5.3.3. Channel Data Type and Channel Order Fields
    • 5.4. Variables
      • 5.4.1. Variable Declarations
      • 5.4.2. Vectors
      • 5.4.3. Array Declarations
      • 5.4.4. Initializers
      • 5.4.5. Alignment
      • 5.4.6. Parameterized Variable Names
      • 5.4.7. Variable Attributes
      • 5.4.8. Variable and Function Attribute Directive: .attribute
    • 5.5. Tensors
      • 5.5.1. Tensor Dimension, size and format
        • 5.5.1.1. Sub-byte Types
          • 5.5.1.1.1. Padding and alignment of the sub-byte types
      • 5.5.2. Tensor Access Modes
      • 5.5.3. Tiled Mode
        • 5.5.3.1. Bounding Box
        • 5.5.3.2. Traversal-Stride
        • 5.5.3.3. Out of Boundary Access
        • 5.5.3.4. .tile::scatter4 and .tile::gather4 modes
          • 5.5.3.4.1. Bounding Box
      • 5.5.4. im2col mode
        • 5.5.4.1. Bounding Box
        • 5.5.4.2. Traversal Stride
        • 5.5.4.3. Out of Boundary Access
      • 5.5.5. im2col::w, im2col_no_offs::w and im2col::w::128 modes
        • 5.5.5.1. Bounding Box
        • 5.5.5.2. Traversal Stride
        • 5.5.5.3. wHalo
        • 5.5.5.4. wOffset
      • 5.5.6. Interleave layout
      • 5.5.7. Swizzling Modes
      • 5.5.8. Tensor-map
  • 6. Instruction Operands
    • 6.1. Operand Type Information
    • 6.2. Source Operands
    • 6.3. Destination Operands
    • 6.4. Using Addresses, Arrays, and Vectors
      • 6.4.1. Addresses as Operands
        • 6.4.1.1. Generic Addressing
      • 6.4.2. Arrays as Operands
      • 6.4.3. Vectors as Operands
      • 6.4.4. Labels and Function Names as Operands
    • 6.5. Type Conversion
      • 6.5.1. Scalar Conversions
      • 6.5.2. Rounding Modifiers
    • 6.6. Operand Costs
  • 7. Abstracting the ABI
    • 7.1. Function Declarations and Definitions
      • 7.1.1. Changes from PTX ISA Version 1.x
    • 7.2. Variadic Functions
    • 7.3. Alloca
  • 8. Memory Consistency Model
    • 8.1. Scope and applicability of the model
      • 8.1.1. Limitations on atomicity at system scope
    • 8.2. Memory operations
      • 8.2.1. Overlap
      • 8.2.2. Aliases
      • 8.2.3. Multimem Addresses
      • 8.2.4. Memory Operations on Vector Data Types
      • 8.2.5. Memory Operations on Packed Data Types
      • 8.2.6. Initialization
    • 8.3. State spaces
    • 8.4. Operation types
      • 8.4.1. mmio Operation
      • 8.4.2. volatile Operation
    • 8.5. Scope
    • 8.6. Proxies
      • 8.6.1. Strong Proxy Accesses
    • 8.7. Morally strong operations
      • 8.7.1. Conflict and Data-races
      • 8.7.2. Limitations on Mixed-size Data-races
    • 8.8. Release and Acquire Patterns
    • 8.9. Ordering of memory operations
      • 8.9.1. Program Order
        • 8.9.1.1. Asynchronous Operations
      • 8.9.2. Observation Order
      • 8.9.3. Fence-SC Order
      • 8.9.4. Memory synchronization
      • 8.9.5. Causality Order
      • 8.9.6. Coherence Order
      • 8.9.7. Communication Order
    • 8.10. Axioms
      • 8.10.1. Coherence
      • 8.10.2. Fence-SC
      • 8.10.3. Atomicity
      • 8.10.4. No Thin Air
      • 8.10.5. Sequential Consistency Per Location
      • 8.10.6. Causality
    • 8.11. Special Cases
      • 8.11.1. Reductions do not form Acquire Patterns
  • 9. Instruction Set
    • 9.1. Format and Semantics of Instruction Descriptions
    • 9.2. PTX Instructions
    • 9.3. Predicated Execution
      • 9.3.1. Comparisons
        • 9.3.1.1. Integer and Bit-Size Comparisons
        • 9.3.1.2. Floating Point Comparisons
      • 9.3.2. Manipulating Predicates
    • 9.4. Type Information for Instructions and Operands
      • 9.4.1. Operand Size Exceeding Instruction-Type Size
    • 9.5. Divergence of Threads in Control Constructs
    • 9.6. Semantics
      • 9.6.1. Machine-Specific Semantics of 16-bit Code
    • 9.7. Instructions
      • 9.7.1. Integer Arithmetic Instructions
        • 9.7.1.1. Integer Arithmetic Instructions: add
        • 9.7.1.2. Integer Arithmetic Instructions: sub
        • 9.7.1.3. Integer Arithmetic Instructions: mul
        • 9.7.1.4. Integer Arithmetic Instructions: mad
        • 9.7.1.5. Integer Arithmetic Instructions: clmad
        • 9.7.1.6. Integer Arithmetic Instructions: mul24
        • 9.7.1.7. Integer Arithmetic Instructions: mad24
        • 9.7.1.8. Integer Arithmetic Instructions: sad
        • 9.7.1.9. Integer Arithmetic Instructions: div
        • 9.7.1.10. Integer Arithmetic Instructions: rem
        • 9.7.1.11. Integer Arithmetic Instructions: abs
        • 9.7.1.12. Integer Arithmetic Instructions: neg
        • 9.7.1.13. Integer Arithmetic Instructions: min
        • 9.7.1.14. Integer Arithmetic Instructions: max
        • 9.7.1.15. Integer Arithmetic Instructions: popc
        • 9.7.1.16. Integer Arithmetic Instructions: clz
        • 9.7.1.17. Integer Arithmetic Instructions: bfind
        • 9.7.1.18. Integer Arithmetic Instructions: fns
        • 9.7.1.19. Integer Arithmetic Instructions: brev
        • 9.7.1.20. Integer Arithmetic Instructions: bfe
        • 9.7.1.21. Integer Arithmetic Instructions: bfi
        • 9.7.1.22. Integer Arithmetic Instructions: szext
        • 9.7.1.23. Integer Arithmetic Instructions: bmsk
        • 9.7.1.24. Integer Arithmetic Instructions: dp4a
        • 9.7.1.25. Integer Arithmetic Instructions: dp2a
      • 9.7.2. Extended-Precision Integer Arithmetic Instructions
        • 9.7.2.1. Extended-Precision Arithmetic Instructions: add.cc
        • 9.7.2.2. Extended-Precision Arithmetic Instructions: addc
        • 9.7.2.3. Extended-Precision Arithmetic Instructions: sub.cc
        • 9.7.2.4. Extended-Precision Arithmetic Instructions: subc
        • 9.7.2.5. Extended-Precision Arithmetic Instructions: mad.cc
        • 9.7.2.6. Extended-Precision Arithmetic Instructions: madc
      • 9.7.3. Floating-Point Instructions
        • 9.7.3.1. Floating Point Instructions: testp
        • 9.7.3.2. Floating Point Instructions: copysign
        • 9.7.3.3. Floating Point Instructions: add
        • 9.7.3.4. Floating Point Instructions: sub
        • 9.7.3.5. Floating Point Instructions: mul
        • 9.7.3.6. Floating Point Instructions: fma