-
Notifications
You must be signed in to change notification settings - Fork 2.5k
All issues
Issue creation is restricted in this repository
Issues
is:issue state:open
is:issue state:open
Search results
[Improvement] Stream pre-combine deduplication for sorted Flink LSM batches
type:devtaskDevelopment tasks and maintenance workDevelopment tasks and maintenance workStatus: Open.#20123 In apache/hudi;[BUG] Native log writer skips delete writer cleanup when data writer close fails
area:writerWrite client and core write operationsWrite client and core write operationstype:bugBug reports and fixesBug reports and fixesStatus: Open.#20117 In apache/hudi;Spark file index returns one PartitionDirectory per file slice instead of one per partition
area:performancePerformance optimizationsPerformance optimizationsarea:readerReader core functionalityReader core functionalitypriority:highSignificant impact; potential bugsSignificant impact; potential bugstype:bugBug reports and fixesBug reports and fixesStatus: Open.#20114 In apache/hudi;A DataFrame write does not invalidate a CACHE TABLE'd Hudi table, so queries silently return stale rows
type:bugBug reports and fixesBug reports and fixesStatus: Open.#20113 In apache/hudi;Add functional guards against executor .hoodie access and task payload growth on the Spark read path
area:performancePerformance optimizationsPerformance optimizationsarea:testsTesting-relatedTesting-relatedpriority:mediumModerate impact; usability gapsModerate impact; usability gapstype:improvementImprovements to existing functionalityImprovements to existing functionalityStatus: Open.#20111 In apache/hudi;Spark row-based parquet reads regenerate the same row projection for every file
area:performancePerformance optimizationsPerformance optimizationsarea:readerReader core functionalityReader core functionalitypriority:mediumModerate impact; usability gapsModerate impact; usability gapstype:improvementImprovements to existing functionalityImprovements to existing functionalityStatus: Open.#20109 In apache/hudi;Schema-on-read batch reads return wrong values after some column type changes
area:readerReader core functionalityReader core functionalityarea:schemaSchema evolution and data typesSchema evolution and data typespriority:highSignificant impact; potential bugsSignificant impact; potential bugstype:bugBug reports and fixesBug reports and fixesStatus: Open.#20108 In apache/hudi;- Status: Open.#20106 In apache/hudi;
CACHE TABLE on a Hudi table strands the cached dataset on every write — fileStatusCache participates in HoodieFileIndex equality
type:bugBug reports and fixesBug reports and fixesStatus: Open.#20104 In apache/hudi;Spark parquet base file reads convert and compare the full footer schema for every file
area:performancePerformance optimizationsPerformance optimizationsarea:readerReader core functionalityReader core functionalitypriority:mediumModerate impact; usability gapsModerate impact; usability gapstype:improvementImprovements to existing functionalityImprovements to existing functionalityStatus: Open.#20101 In apache/hudi;Spark parquet base file reads repeat conf copies and schema work that is constant for the scan
area:performancePerformance optimizationsPerformance optimizationsarea:readerReader core functionalityReader core functionalitypriority:mediumModerate impact; usability gapsModerate impact; usability gapstype:improvementImprovements to existing functionalityImprovements to existing functionalityStatus: Open.#20100 In apache/hudi;Table service tasks ship the table, and bootstrap listing ignores the job's Hadoop configuration
area:performancePerformance optimizationsPerformance optimizationsarea:table-serviceTable servicesTable servicespriority:mediumModerate impact; usability gapsModerate impact; usability gapstype:improvementImprovements to existing functionalityImprovements to existing functionalityStatus: Open.#20094 In apache/hudi;