Ryujinx

mirror of https://git.naxdy.org/Mirror/Ryujinx.git synced 2025-02-18 15:23:36 +00:00

Author	SHA1	Message	Date
gdkchan	43ebd7a9bb	New shader cache implementation (#3194 ) * New shader cache implementation * Remove some debug code * Take transform feedback varying count into account * Create shader cache directory if it does not exist + fragment output map related fixes * Remove debug code * Only check texture descriptors if the constant buffer is bound * Also check CPU VA on GetSpanMapped * Remove more unused code and move cache related code * XML docs + remove more unused methods * Better codegen for TransformFeedbackDescriptor.AsSpan * Support migration from old cache format, remove more unused code Shader cache rebuild now also rewrites the shared toc and data files * Fix migration error with BRX shaders * Add a limit to the async translation queue Avoid async translation threads not being able to keep up and the queue growing very large * Re-create specialization state on recompile This might be required if a new version of the shader translator requires more or less state, or if there is a bug related to the GPU state access * Make shader cache more error resilient * Add some missing XML docs and move GpuAccessor docs to the interface/use inheritdoc * Address early PR feedback * Fix rebase * Remove IRenderer.CompileShader and IShader interface, replace with new ShaderSource struct passed to CreateProgram directly * Handle some missing exceptions * Make shader cache purge delete both old and new shader caches * Register textures on new specialization state * Translate and compile shaders in forward order (eliminates diffs due to different binding numbers) * Limit in-flight shader compilation to the maximum number of compilation threads * Replace ParallelDiskCacheLoader state changed event with a callback function * Better handling for invalid constant buffer 1 data length * Do not create the old cache directory structure if the old cache does not exist * Constant buffer use should be per-stage. This change will invalidate existing new caches (file format version was incremented) * Replace rectangle texture with just coordinate normalization * Skip incompatible shaders that are missing texture information, instead of crashing This is required if we, for example, support new texture instruction to the shader translator, and then they allow access to textures that were not accessed before. In this scenario, the old cache entry is no longer usable * Fix coordinates normalization on cubemap textures * Check if title ID is null before combining shader cache path * More robust constant buffer address validation on spec state * More robust constant buffer address validation on spec state (2) * Regenerate shader cache with one stream, rather than one per shader. * Only create shader cache directory during initialization * Logging improvements * Proper shader program disposal * PR feedback, and add a comment on serialized structs * XML docs for RegisterTexture Co-authored-by: riperiperi <rhy3756547@hotmail.com>	2022-04-10 10:49:44 -03:00
gdkchan	e44a43c7e1	Implement VMAD shader instruction and improve InvocationInfo and ISBERD handling (#3251 ) * Implement VMAD shader instruction and improve InvocationInfo and ISBERD handling * Shader cache version bump * Fix typo	2022-04-08 12:42:39 +02:00
gdkchan	3139a85a2b	Allow copy texture views to have mismatching multisample state (#3152 )	2022-04-08 11:26:48 +02:00
merry	a4e8bea866	Lop3Expression: Optimize expressions (#3184 ) * lut3 * bugfixes * TruthTable * false/true -> 0/-1 * add or to expressions * fix inversions * increment cache version	2022-04-08 11:17:38 +02:00
gdkchan	952f6f8a65	Calculate vertex buffer size from index buffer type (#3253 ) * Calculate vertex buffer size from index buffer type * We also need to update the size if first vertex changes	2022-04-08 11:02:06 +02:00
gdkchan	d4b960d348	Implement primitive restart draw arrays properly on OpenGL (#3256 )	2022-04-04 18:43:24 -03:00
gdkchan	b2a225558d	Do not force scissor on clear if scissor is disabled (#3258 )	2022-04-04 18:30:43 -03:00
gdkchan	1402d8391d	Support NVDEC H264 interlaced video decoding and VIC deinterlacing (#3225 ) * Support NVDEC H264 interlaced video decoding and VIC deinterlacing * Remove unused code	2022-03-23 17:09:32 -03:00
gdkchan	79408b68c3	De-tile GOB when DMA copying from block linear to pitch kind memory regions (#3207 ) * De-tile GOB when DMA copying from block linear to pitch kind memory regions * XML docs + nits * Remove using * No flush for regular buffer copies * Add back ulong casts, fix regression due to oversight	2022-03-20 13:55:07 -03:00
gdkchan	e5ad1dfa48	Implement S8D24 texture format and tweak depth range detection (#2458 )	2022-03-15 03:42:08 +01:00
gdkchan	79becc4b78	Dynamically increase buffer size when resizing (#2861 ) * Grow buffers by 1.5x of its size when resizing * Further restrict the cases where the dynamic expansion is done	2022-03-15 03:33:53 +01:00
gdkchan	0bcbe32367	Only initialize shader outputs that are actually used on the next stage (#3054 ) * Only initialize shader outputs that are actually used on the next stage * Shader cache version bump	2022-03-06 20:42:13 +01:00
gdkchan	0a24aa6af2	Allow textures to have their data partially mapped (#2629 ) * Allow textures to have their data partially mapped * Explicitly check for invalid memory ranges on the MultiRangeList * Update GetWritableRegion to also support unmapped ranges	2022-02-22 13:34:16 -03:00
riperiperi	c9c65af59e	Perform unscaled 2d engine copy on CPU if source texture isn't in cache. (#3112 ) * Initial implementation of fast 2d copy TODO: Partial copy for mismatching region/size. * WIP * Cleanup * Update Ryujinx.Graphics.Gpu/Engine/Twod/TwodClass.cs Co-authored-by: gdkchan <gab.dark.100@gmail.com> Co-authored-by: gdkchan <gab.dark.100@gmail.com>	2022-02-22 11:21:29 -03:00
Berkan Diler	644b497df1	Collapse AsSpan().Slice(..) calls into AsSpan(..) (#3145 ) * Collapse AsSpan().Slice(..) calls into AsSpan(..) Less code and a bit faster * Collapse an Array.Clear(array, 0, array.Length) call to Array.Clear(array)	2022-02-22 10:32:10 -03:00
gdkchan	72e543e946	Prefer texture over textureSize for sampler type (#3132 ) * Prefer texture over textureSize for sampler type * Shader cache version bump	2022-02-18 02:44:46 +01:00
gdkchan	3bd357045f	Do not allow render targets not explicitly written by the fragment shader to be modified (#3063 ) * Do not allow render targets not explicitly written by the fragment shader to be modified * Shader cache version bump * Remove blank lines * Avoid redundant color mask updates * HostShaderCacheEntry can be null * Avoid more redundant glColorMask calls * nit: Mask -> Masks * Fix currentComponentMask * More efficient way to update _currentComponentMasks	2022-02-16 23:15:39 +01:00
gdkchan	7bfb5f79b8	When copying linear textures, DMA should ignore region X/Y (#3121 )	2022-02-16 11:13:45 +01:00
Berkan Diler	8f35345729	Use Enum and Delegate.CreateDelegate generic overloads (#3111 ) * Use Enum generic overloads * Remove EnumExtensions.cs * Use Delegate.CreateDelegate generic overloads	2022-02-13 10:50:07 -03:00
gdkchan	f861f0bca2	Fix missing geometry shader passthrough inputs (#3106 ) * Fix missing geometry shader passthrough inputs * Shader cache version bump	2022-02-11 19:52:20 +01:00
Mary	6dffe0fad4	misc: Make PID unsigned long instead of long (#3043 )	2022-02-09 17:18:07 -03:00
gdkchan	b944941733	Fix bug that could cause depth buffer to be missing after clear (#3067 )	2022-01-31 00:11:43 -03:00
riperiperi	c52158b733	Add timestamp to 16-byte/4-word semaphore releases. (#3049 ) * Add timestamp to 16-byte semaphore releases. BOTW was reading a ulong 8 bytes after a semaphore return. Turns out this is the timestamp it was trying to do performance calculation with, so I've made it write when necessary. This mode was also added to the DMA semaphore I added recently, as it is required by a few games. (i think quake?) The timestamp code has been moved to GPU context. Check other games with an unusually low framerate cap or dynamic resolution to see if they have improved. * Cast dma semaphore payload to ulong to fill the space * Write timestamp first Might be just worrying too much, but we don't want the applcation reading timestamp if it sees the payload before timestamp is written.	2022-01-27 22:50:32 +01:00
riperiperi	fd6d3ec88f	Fix res scale parameters not being updated in vertex shader (#3046 ) This fixes an issue where the render scale array would not be updated when technically the scales on the flat array were the same, but the start index for the vertex scales was different.	2022-01-27 14:17:13 -03:00
gdkchan	42c75dbb8f	Add support for BC1/2/3 decompression (for 3D textures) (#2987 ) * Add support for BC1/2/3 decompression (for 3D textures) * Optimize and clean up * Unsafe not needed here * Fix alpha value interpolation when a0 <= a1	2022-01-22 19:23:00 +01:00
gdkchan	7e967d796c	Stop using glTransformFeedbackVaryings and use explicit layout on the shader (#3012 ) * Stop using glTransformFeedbackVarying and use explicit layout on the shader * This is no longer needed * Shader cache version bump * Fix gl_PerVertex output for tessellation control shaders	2022-01-21 12:35:21 -03:00
gdkchan	0e59573f2b	Add capability for BGRA formats (#3011 )	2022-01-20 08:37:21 -03:00
gdkchan	fb853f13e9	Scale scissor used for clears (#3002 )	2022-01-16 20:23:00 -03:00
gdkchan	6e0799580f	Fix render target clear when sizes mismatch (#2994 )	2022-01-11 20:15:17 +01:00
riperiperi	ef24c8983d	Fix adjacent 3d texture slices being detected as Incompatible Overlaps (#2993 ) This fixes some regressions caused by #2971 which caused rendered 3D texture data to be lost for most slices. Fixes issues with Xenoblade 2's colour grading, probably a ton of other games. This also removes the check from TextureCache, making it the tiniest bit smaller (any win is a win here).	2022-01-11 09:37:40 +01:00
gdkchan	7f6b3d234a	Implement IMUL, PCNT and CONT shader instructions, fix FFMA32I and HFMA32I (#2972 ) * Implement IMUL shader instruction * Implement PCNT/CONT instruction and fix FFMA32I * Add HFMA232I to the table * Shader cache version bump * No Rc on Ffma32i	2022-01-10 12:08:00 -03:00
gdkchan	952c6e4d45	Fix sampled multisample image size (#2984 )	2022-01-10 08:45:25 +01:00
riperiperi	cda659955c	Texture Sync, incompatible overlap handling, data flush improvements. (#2971 ) * Initial test for texture sync * WIP new texture flushing setup * Improve rules for incompatible overlaps Fixes a lot of issues with Unreal Engine games. Still a few minor issues (some caused by dma fast path?) Needs docs and cleanup. * Cleanup, improvements Improve rules for fast DMA * Small tweak to group together flushes of overlapping handles. * Fixes, flush overlapping texture data for ASTC and BC4/5 compressed textures. Fixes the new Life is Strange game. * Flush overlaps before init data, fix 3d texture size/overlap stuff * Fix 3D Textures, faster single layer flush Note: nosy people can no longer merge this with Vulkan. (unless they are nosy enough to implement the new backend methods) * Remove unused method * Minor cleanup * More cleanup * Use the More Fun and Hopefully No Driver Bugs method for getting compressed tex too This one's for metro * Address feedback, ASTC+ETC to FormatClass * Change offset to use Span slice rather than IntPtr Add * Fix this too	2022-01-09 13:28:48 -03:00
riperiperi	79adba4402	Add support for render scale to vertex stage. (#2763 ) * Add support for render scale to vertex stage. Occasionally games read off textureSize on the vertex stage to inform the fragment shader what size a texture is without querying in there. Scales were not present in the vertex shader to correct the sizes, so games were providing the raw upscaled texture size to the fragment shader, which was incorrect. One downside is that the fragment and vertex support buffer description must be identical, so the full size scales array must be defined when used. I don't think this will have an impact though. Another is that the fragment texture count must be updated when vertex shader textures are used. I'd like to correct this so that the update is folded into the update for the scales. Also cleans up a bunch of things, like it making no sense to call CommitRenderScale for each stage. Fixes render scale causing a weird offset bloom in Super Mario Party and Clubhouse Games. Clubhouse Games still has a pixelated look in a number of its games due to something else it does in the shader. * Split out support buffer update, lazy updates. * Commit support buffer before compute dispatch * Remove unnecessary qualifier. * Address Feedback	2022-01-08 14:48:48 -03:00
gdkchan	15131d4350	Force crop when presentation cached texture size mismatches (#2957 )	2021-12-31 12:00:42 -03:00
gdkchan	c05c8e09d4	Add support for the R4G4 texture format (#2956 )	2021-12-30 17:10:54 +01:00
gdkchan	ef39b2ebdd	Flip scissor box when the YNegate bit is set (#2941 ) * Flip scissor box when the YNegate bit is set * Flip scissor based on screen scissor state, account for negative scissor Y * No need for abs when we already know the value is negative	2021-12-28 08:37:23 -03:00
gdkchan	a87f7f2029	Fix DMA copy fast path line size when xCount < stride (#2942 )	2021-12-26 13:05:26 -03:00
gdkchan	451673ada5	Fix I2M texture copies when line length is not a multiple of 4 (#2938 ) * Fix I2M texture copies when line length is not a multiple of 4 * Do not copy padding bytes for 1D copies * Nit	2021-12-26 12:39:07 -03:00
gdkchan	e7c2dc8ec3	Fix for texture pool not being updated when it should + buffer texture related fixes (#2911 )	2021-12-19 11:50:44 -03:00
riperiperi	521a07e612	Add support for releasing a semaphore to DmaClass (#2926 ) * Add support for releasing a semaphore to DmaClass Fixes freezes in OpenGL games, primarily GameMaker ones such as Undertale. * Address Feedback	2021-12-19 11:32:52 -03:00
gdkchan	119a3a1887	Fix SUATOM and other texture shader instructions with RZ dest (#2885 ) * Fix SUATOM and other texture shader instructions with RZ dest * Shader cache version bump	2021-12-08 18:36:09 -03:00
riperiperi	bc4e70b6fa	Move texture anisotropy check to SetInfo (#2843 ) Rather than calculating this for every sampler, this PR calculates if a texture can force anisotropy when its info is set, and exposes the value via a public boolean. This should help texture/sampler heavy games when anisotropic filtering is not Auto, like UE4 ones (or so i hear?). There is another cost where samplers are created twice when anisotropic filtering is enabled, but I'm not sure how relevant this one is.	2021-12-08 18:09:36 -03:00
gdkchan	650cc41c02	Implement remaining shader double-precision instructions (#2845 ) * Implement remaining shader double-precision instructions * Shader cache version bump	2021-12-08 17:54:12 -03:00
gdkchan	acc0b0f313	Fix FLO.SH shader instruction with a input of 0 (#2876 ) * Fix FLO.SH shader instruction with a input of 0 * Shader cache version bump	2021-12-05 13:25:05 +01:00
Mary	57d3296ba4	infra: Migrate to .NET 6 (#2829 ) * infra: Migrate to .NET 6 * Rollback version naming change * Workaround .NET 6 ZipArchive API issues * ci: Switch to VS 2022 for AppVeyor CI is now ready for .NET 6 * Suppress WebClient warning in DoUpdateWithMultipleThreads * Attempt to workaround System.Drawing.Common changes on 6.0.0 * Change keyboard rendering from System.Drawing to ImageSharp * Make the software keyboard renderer multithreaded * Bump ImageSharp version to 1.0.4 to fix a bug in Image.Load * Add fallback fonts to the keyboard renderer * Fix warnings * Address caian's comment * Clean up linux workaround as it's uneeded now * Update readme Co-authored-by: Caian Benedicto <caianbene@gmail.com>	2021-11-28 21:24:17 +01:00
gdkchan	30b7aaefca	Better depth range detection (#2754 ) * Better depth range detection * PR feedback * Move depth mode set out of the loop and to a separate method	2021-11-21 10:25:03 -03:00
riperiperi	788aec511f	Limit Custom Anisotropic Filtering to mipmapped textures with many levels (#2832 ) * Limit Custom Anisotropic Filtering to only fully mipmapped textures There's a major flaw with the anisotropic filtering setting that causes @GamerzHell9137 to report graphical bugs that otherwise wouldn't be there, because he just won't set it to Auto. This should fix those issues, hopefully. These bugs are generally because anisotropic filtering is enabled on something that it shouldn't be, such as a post process filter or some data texture. This PR maintains two host samplers when custom AF is enabled, and only uses the forced AF one when the texture is 2d and fully mipmapped (goes down to 1x1). This is because game textures are the ideal target for this filtering, and they are typically fully mipmapped, unlike things like screen render targets which usually have 1 or just a few levels. This also only enables AF on mipmapped samplers where the filtering is bilinear or trilinear. This should be self explanatory. This PR also allows the changing of Anisotropic Filtering at runtime, and you can immediately see the changes. All samplers are flushed from the cache if the setting changes, causing them to be recreated with the new custom AF value. This brings it in line with our resolution scale. 😌 * Expected minimum mip count for large textures rather than all, address feedback * Use Target rather than Info.Target * Retrigger build? * Fix rebase	2021-11-13 16:04:21 -03:00
gdkchan	611bec6e44	Implement DrawTexture functionality (#2747 ) * Implement DrawTexture functionality * Non-NVIDIA support * Disable some features that should not affect draw texture (slow path) * Remove space from shader source * Match 2D engine names * Fix resolution scale and add missing XML docs * Disable transform feedback for draw texture fallback	2021-11-10 15:37:49 -03:00
gdkchan	911ea38e93	Support shader gl_Color, gl_SecondaryColor and gl_TexCoord built-ins (#2817 ) * Support shader gl_Color, gl_SecondaryColor and gl_TexCoord built-ins * Shader cache version bump * Fix back color value on fragment shader * Disable IPA multiplication for fixed function attributes and back color selection	2021-11-08 13:18:46 -03:00
gdkchan	3dee712164	Fix bindless/global memory elimination with inverted predicates (#2826 ) * Fix bindless/global memory elimination with inverted predicates * Shader cache version bump	2021-11-08 12:57:28 -03:00
gdkchan	b7a1544e8b	Fix InvocationInfo on geometry shader and bindless default integer const (#2822 ) * Fix InvocationInfo on geometry shader and bindless default integer const * Shader cache version bump * Consistency for the default value	2021-11-08 11:39:30 -03:00
gdkchan	e48530e9d9	When waiting on CPU, do not return a time out error from EventWait (#2780 ) * When waiting on CPU, do not return a time out error from EventWait * And while I'm at it...	2021-11-01 19:10:02 -03:00
gdkchan	99445dd0a6	Add support for fragment shader interlock (#2768 ) * Support coherent images * Add support for fragment shader interlock * Change to tree based match approach * Refactor + check for branch targets and external registers * Make detection more robust * Use Intel fragment shader ordering if interlock is not available, use nothing if both are not available * Remove unused field	2021-10-28 19:53:12 -03:00
gdkchan	0d174cbd45	EventWait should not signal the event when it returns Success (#2739 ) * Fix race when EventWait is called and a wait is done on the CPU * This is useless now * Fix EventSignal * Ensure the signal belongs to the current fence, to avoid stale signals	2021-10-19 17:25:32 -03:00
gdkchan	63f1663fa9	Fix shader 8-bit and 16-bit STS/STG (#2741 ) * Fix 8 and 16-bit STG * Fix 8 and 16-bit STS * Shader cache version bump	2021-10-18 20:24:15 -03:00
riperiperi	052deebf26	Another workaround for NVIDIA driver 496.13 shader bug (#2750 ) * Another workaround for NVIDIA driver 496.13 shader bug This might work better than the other one. Give this a test to see if it fixes/doesn't fix issues with the other one. The problem seems to be when any variable assignment happens with a negation. `temp_1 = -temp_0;` seems to trigger weird behaviour, but `temp_1 = 0.0 - temp_0;` does not. This also might to extend towards integer types? * Update cache version * Add disclaimer comment * Wording	2021-10-18 20:04:06 -03:00
gdkchan	d512ce122c	Initial tessellation shader support (#2534 ) * Initial tessellation shader support * Nits * Re-arrange built-in table * This is not needed anymore * PR feedback	2021-10-18 18:38:04 -03:00
gdkchan	25fd4ef10e	Extend bindless elimination to work with masked and shifted handles (#2727 ) * Extent bindless elimination to work with masked handles * Extend bindless elimination to catch shifted pattern, refactor handle packing/unpacking	2021-10-17 17:28:18 -03:00
gdkchan	d05573bfd1	Implement SHF (funnel shift) shader instruction (#2702 ) * Implement SHF shader instruction * Shader cache version bump * Better name	2021-10-17 17:02:20 -03:00
gdkchan	464a92d8a7	Force index buffer update for games using Vulkan (#2726 )	2021-10-12 23:46:42 +02:00
riperiperi	0bce4a074a	Don't force scaling on 2D copy sources (#2701 ) Some games (GameMaker Studio) build texture atlases out of sprites during initialization, using the 2D copy method. These copies are done from textures loaded into memory, not rendered, so they are not scaled to begin with. I had set srcTexture in these copies to force scaling, but really it only needs to scale if the texture already exists and was scaled by rendering or something else. I just set that to false, so it doesn't change if the texture is scaled or not. This will also avoid the destination being scaled if the source wasn't. The copy can handle mismatching scales just fine. This prevents scaling artifacts in GMS games, and maybe others (not Super Mario Maker 2, that has another issue).	2021-10-12 23:12:17 +02:00
gdkchan	a7109c767b	Rewrite shader decoding stage (#2698 ) * Rewrite shader decoding stage * Fix P2R constant buffer encoding * Fix PSET/PSETP * PR feedback * Log unimplemented shader instructions * Implement NOP * Remove using * PR feedback	2021-10-12 22:35:31 +02:00
riperiperi	a4956591ec	Avoid potential race	2021-10-07 01:13:51 +01:00
riperiperi	c61c1ea898	Reregister flush actions when taking a buffer's modified range list. Fixes a regression from #2663 where buffer flush would not happen after a resize. Specifically caused the world map in Yoshi's Crafted World to flash. I have other planned changes to this class so this might change soon, but this regression could affect a lot so it couldn't wait.	2021-10-07 00:00:56 +01:00
riperiperi	fff48bb45a	Smaller initial size for ModifiedRangeList & directly inherit range list (#2663 ) This fixes a potential regression with the new range list changes, where the cost for creating new ones would be rather large due to creating a 1024 size array. Also reduces cost for range list inheritance by using the first existing range list as a base, rather than creating a new one then adding both lists to it. The growth size for the RangeList is now identical to its initial size. Every 32 elements was probably a little too common - now it is 1024 for most things and 8 for the buffer modified range list. The Unmapped and SyncMethod methods have been changed to ensure that they behave properly if the range list is set null. Cleaned up a few calls to use the null-conditional operator.	2021-10-04 15:38:59 -03:00
gdkchan	75f4b1ff2d	Relax sampler pool requirement (#2703 )	2021-10-04 14:35:28 -03:00
riperiperi	d92fff541b	Replace CacheResourceWrite with more general "precise" write (#2684 ) * Replace CacheResourceWrite with more general "precise" write The goal of CacheResourceWrite was to notify GPU resources when they were modified directly, by looking up the modified address/size in a structure and calling a method on each resource. The downside of this is that each resource cache has to be queried individually, they all have to implement their own way to do this, and it can only signal to resources using the same PhysicalMemory instance. This PR adds the ability to signal a write as "precise" on the tracking, which signals a special handler (if present) which can be used to avoid unnecessary flush actions, or maybe even more. For buffers, precise writes specifically do not flush, and instead punch a hole in the modified range list to indicate that the data on GPU has been replaced. The downside is that precise actions must ignore the page protection bits and always signal - as they need to notify the target resource to ignore the sequence number optimization. I had to reintroduce the sequence number increment after I2M, as removing it was causing issues in rabbids kingdom battle. However - all resources modified by I2M are notified directly to lower their sequence number, so the problem is likely that another unrelated resource is not being properly updated. Thankfully, doing this does not affect performance in the games I tested. This should fix regressions from #2624. Test any games that were broken by that. (RF4, rabbids kingdom battle) I've also added a sequence number increment to ThreedClass.IncrementSyncpoint, as it seems to fix buffer corruption in OpenGL homebrew. (this was a regression from removing sequence number increment from constant buffer update - another unrelated resource thing) * Add tests. * Add XML docs for GpuRegionHandle * Skip UpdateProtection if only precise actions were called This allows precise actions to skip reprotection costs.	2021-09-29 02:27:03 +02:00
riperiperi	b6e093b0fc	Force copy when auto-deleting a texture with dependencies (#2687 ) When a texture is deleted by falling to the bottom of the AutoDeleteCache, its data is flushed to preserve any GPU writes that occurred. This ensures that the data appears in any textures recreated in the future, but didn't account for a texture that already existed with a copy dependency. This change forces copy dependencies to complete if a texture falls out from from the AutoDeleteCache. (not removed via overlap, as that would be wasted effort) Fixes broken lighting caused by pausing in SMO's Metro Kingdom. May fix some other issues.	2021-09-29 02:11:05 +02:00
gdkchan	fd7567a6b5	Only make render target 2D textures layered if needed (#2646 ) * Only make render target 2D textures layered if needed * Shader cache version bump * Ensure topology is updated on channel swap	2021-09-29 01:55:12 +02:00
gdkchan	83bdafccda	Share scales array for graphics and compute (#2653 )	2021-09-28 23:52:27 +02:00
riperiperi	7c5ead1c19	Fast path for Inline2Memory buffer write that skips write tracking (#2624 ) * Fast path for Inline2Memory buffer write This PR adds a method to PhysicalMemory that attempts to write all cached resources directly, so that memory tracking can be avoided. The goal of this is both to avoid flushing buffer data, and to avoid raising the sequence number when data is written, which causes buffer and texture handles to be re-checked. This currently only targets buffers, with a side check on textures that falls back to a tracked write if any exist within the target range. It's not expected to write textures from here - this is just a mechanism to protect us if someone does decide to do that. It's possible to add a fast path for this in future (and for ShaderCache, once that starts using tracking) The forced read before inline2memory begins has been skipped, as the data is fully written when the transfer is completed anyways. This allows us to flush on read in emergency situations, but still write the new data over the flushed data. Improves performance on Xenoblade 2 and DE, which was flushing buffer data on the GPU thread when trying to write compute data. May improve performance in other games that write SSBOs from compute, and update data in the same/nearby pages often. Super Smash Bros Ultimate should probably be tested to make sure the vertex explosions haven't returned, as I think that's what this AdvanceSequence was for. * ForceDirty before write, to make sure data does not flush over the new write	2021-09-19 15:09:53 +02:00
gdkchan	f08a280ade	Use shader subgroup extensions if shader ballot is not supported (#2627 ) * Use shader subgroup extensions if shader ballot is not supported * Shader cache version bump + cleanup * The type is still required on the table	2021-09-19 14:38:39 +02:00
riperiperi	7379bc2f39	Array based RangeList that caches Address/EndAddress (#2642 ) * Array based RangeList that caches Address/EndAddress In isolation, this was more than 2x faster than the RangeList that checks using the interface. In practice I'm seeing much better results than I expected. The array is used because checking it is slightly faster than using a list, which loses time to struct copies, but I still want that data locality. A method has been added to the list to update the cached end address, as some users of the RangeList currently modify it dynamically. Greatly improves performance in Super Mario Odyssey, Xenoblade and any other GPU limited games. * Address Feedback	2021-09-19 14:22:26 +02:00
riperiperi	b0af010247	Set texture/image bindings in place rather than allocating and passing an array (#2647 ) * Remove allocations for texture bindings and state * Rent rather than stackalloc + copy A bit faster.	2021-09-19 14:03:05 +02:00
gdkchan	ac4ec1a015	Account for negative strides on DMA copy (#2623 ) * Account for negative strides on DMA copy * Should account for non-zero Y	2021-09-11 22:54:18 +02:00
riperiperi	b0e410a828	Lift textures in the AutoDeleteCache for all modifications. (#2615 ) * Lift textures in the AutoDeleteCache for all modifications. Before, this would only apply to render targets and texture blit. Now it applies to image stores, the fast dma copy path and any other type of modification. Image store always at least has one reference in the texture pool, so the function of the AutoDeleteCache keeping textures _alive_ is not useful, but a very important function for a while has been its use to flush textures in order of modification when they are dereferenced, so that their data is not lost. Before, textures populated using image stores were being dereferenced and reloaded as garbage. Now, when these textures are dereferenced, their data will be put back into memory, and everything stays intact. Fixes lighting breaking when switching levels in THPS1+2, and potentially some more UE4 games. I've tested a bunch more games for regressions and performance impact, but they all seem fine. * Lift copy srcTexture so that it doesn't remain referenceless * Perform lift before reference count change on unbind. It's important to lift on unbind as that is the moment the texture was truly last modified, but definitely not after releasing every single reference.	2021-09-11 21:52:54 +02:00
riperiperi	f0b00c1ae9	Fix TXQ for 3D textures. (#2613 ) * Fix TXQ for 3D textures. Assumes the texture is 3D if the component mask contains Z. This fixes a bug in UE4 games where parts of the map had garbage pointers to lighting voxels, as the lookup 3D texture was not being initialized. Most notable game is THPS1+2. May need another PR to keep image store data alive and properly flush it in order using the AutoDeleteCache. * Get sampler type for TextureSize from bound textures.	2021-09-02 00:17:43 -03:00
riperiperi	142cededd4	Implement Shader Instructions SUATOM and SURED (#2090 ) * Initial Implementation * Further improvements (no support for float/64-bit types) * Merge atomic and reduce instructions, add missing format switch * Fix rebase issues. * Not used. * Whoops. Fixed. * Partial implementation of inc/dec, cleanup and TODOs * Remove testing path * Address Feedback	2021-08-31 02:51:57 -03:00
gdkchan	416dc8fde4	Fix out-of-bounds shader thread shuffle (#2605 ) * Fix out-of-bounds shader thread shuffle * Shader cache version bump	2021-08-30 14:02:40 -03:00
gdkchan	82cefc8dd3	Handle indirect draw counts with non-zero draw starts properly (#2593 )	2021-08-29 16:52:38 -03:00
riperiperi	15e7fe3ac9	Avoid deleting textures when their data does not overlap. (#2601 ) * Avoid deleting textures when their data does not overlap. It's possible that while two textures start and end addresses indicate an overlap, that the actual data contained within them is sparse due to a layer stride. One such possibility is array slices of a cubemap at different mip levels - they overlap on a whole, but the actual texture data fills the gaps between each other's layers rather than actually overlapping. This fixes issues with UE4 games having incorrect lighting (solid white screen or really dark shadows). There are still remaining issues with games that use the 3D texture prebaked lighting, such as THPS1+2. This PR also fixes a bug with TexturePool's resized texture handling where the base level in the descriptor was not considered. * AllRegions granularity for 3d textures is now by level rather than by slice. * Address feedback	2021-08-29 16:22:13 -03:00
riperiperi	76e8f9ac87	Only reupload the texture scale array if it changes. (#2595 ) * Only reupload the texture scale array if it changes. Before, this would be called all the time if any shader needed a scale value. The cost of doing this has increased with threaded-gal, as the scale array is copied to a span pool, and it's was called on pretty much every draw sometimes. This improves GPU performance in games, scaled or not. Most affected game seems to be Xenoblade Chronicles: Definitive Edition. * Just use = instead of \|=	2021-08-27 17:08:30 -03:00
gdkchan	ee1038e542	Initial support for shader attribute indexing (#2546 ) * Initial support for shader attribute indexing * Support output indexing too, other improvements * Fix order * Address feedback	2021-08-27 01:44:47 +02:00
riperiperi	ec3e848d79	Add a Multithreading layer for the GAL, multi-thread shader compilation at runtime (#2501 ) * Initial Implementation About as fast as nvidia GL multithreading, can be improved with faster command queuing. * Struct based command list Speeds up a bit. Still a lot of time lost to resource copy. * Do shader init while the render thread is active. * Introduce circular span pool V1 Ideally should be able to use structs instead of references for storing these spans on commands. Will try that next. * Refactor SpanRef some more Use a struct to represent SpanRef, rather than a reference. * Flush buffers on background thread * Use a span for UpdateRenderScale. Much faster than copying the array. * Calculate command size using reflection * WIP parallel shaders * Some minor optimisation * Only 2 max refs per command now. The command with 3 refs is gone. 😌 * Don't cast on the GPU side * Remove redundant casts, force sync on window present * Fix Shader Cache * Fix host shader save. * Fixup to work with new renderer stuff * Make command Run static, use array of delegates as lookup Profile says this takes less time than the previous way. * Bring up to date * Add settings toggle. Fix Muiltithreading Off mode. * Fix warning. * Release tracking lock for flushes * Fix Conditional Render fast path with threaded gal * Make handle iteration safe when releasing the lock This is mostly temporary. * Attempt to set backend threading on driver Only really works on nvidia before launching a game. * Fix race condition with BufferModifiedRangeList, exceptions in tracking actions * Update buffer set commands * Some cleanup * Only use stutter workaround when using opengl renderer non-threaded * Add host-conditional reservation of counter events There has always been the possibility that conditional rendering could use a query object just as it is disposed by the counter queue. This change makes it so that when the host decides to use host conditional rendering, the query object is reserved so that it cannot be deleted. Counter events can optionally start reserved, as the threaded implementation can reserve them before the backend creates them, and there would otherwise be a short amount of time where the counter queue could dispose the event before a call to reserve it could be made. * Address Feedback * Make counter flush tracked again. Hopefully does not cause any issues this time. * Wait for FlushTo on the main queue thread. Currently assumes only one thread will want to FlushTo (in this case, the GPU thread) * Add SDL2 headless integration * Add HLE macro commands. Co-authored-by: Mary <mary@mary.zone>	2021-08-27 00:31:29 +02:00
mpnico	8e1adb95cf	Add support for HLE macros and accelerate MultiDrawElementsIndirectCount #2 (#2557 ) * Add support for HLE macros and accelerate MultiDrawElementsIndirectCount * Add missing barrier * Fix index buffer count * Add support check for each macro hle before use * Add missing xml doc Co-authored-by: gdkchan <gab.dark.100@gmail.com>	2021-08-26 23:50:28 +02:00
riperiperi	bdc1f91a5b	Remove pool cache entries for incompatible overlapping textures (#2568 ) This greatly reduces memory usage in games that aggressively reuse memory without removing dead textures from the pool, such as the Xenoblade games, UE3 games, and to a lesser extent, UE4/unity games. This change stops memory usage from ballooning in xenoblade and some other games. It will also reduce texture view/dependency complexity in some games - for example in MK8D it will reduce the number of surface copies between lighting cubemaps generated for actors. There shouldn't be any performance impact from doing this, though the deletion and creation of textures could be improved by improving the OpenGL texture storage cache, which is very simple and limited right now. This will be improved in future. Another potential error has been fixed with the texture cache, which could prevent data loss when data is interchangably written to textures from both the GPU and CPU. It was possible that the dirty flag for a texture would be consumed without the data being synchronized on next use, due to the old overlap check. This check no longer consumes the dirty flag. Please test a bunch of games to make sure they still work, and there are no performance regressions.	2021-08-20 17:52:09 -03:00
riperiperi	97aedc030d	Fix GetHandleInformation for mipmapped 3d textures (#2569 ) Got this the wrong way round - was causing games to try synchronize mipmap levels of like 52 on a 3d texture with 6 levels. Also, corrected the variable name in the method that _was_ working.	2021-08-20 14:59:39 -03:00
gdkchan	680d3ed198	Enable transform feedback buffer flush (#2552 )	2021-08-17 14:09:27 -03:00
gdkchan	eb181425b1	Fix size of cached compute shaders (#2548 ) * Fix size of cached compute shaders * Missed one	2021-08-12 15:59:24 -03:00
gdkchan	8196086f7a	Revert "Calculate vertex buffer sizes from index buffer (#1663 )" (#2544 ) This reverts commit `10d649e6d3`.	2021-08-11 22:13:48 -03:00
gdkchan	3148c0c21c	Unify GpuAccessorBase and TextureDescriptorCapableGpuAccessor (#2542 ) * Unify GpuAccessorBase and TextureDescriptorCapableGpuAccessor * Shader cache version bump	2021-08-11 18:56:59 -03:00
gdkchan	c3e2646f9e	Workaround for Intel FrontFacing built-in variable bug (#2540 )	2021-08-11 23:01:06 +02:00
riperiperi	0a80a837cb	Use "Undesired" scale mode for certain textures rather than blacklisting (#2537 ) * Use "Undesired" scale mode for certain textures rather than blacklisting * Nit Co-authored-by: gdkchan <gab.dark.100@gmail.com> Co-authored-by: gdkchan <gab.dark.100@gmail.com>	2021-08-11 22:44:51 +02:00
gdkchan	ed754af8d5	Make sure attributes used on subsequent shader stages are initialized (#2538 )	2021-08-11 22:27:00 +02:00
gdkchan	10d649e6d3	Calculate vertex buffer sizes from index buffer (#1663 ) * Calculate vertex buffer size from maximum index buffer index * Increase maximum index buffer count for it to be considered profitable for counting	2021-08-11 22:06:09 +02:00
gdkchan	0f6ec446ea	Replace BGRA and scale uniforms with a uniform block (#2496 ) * Replace BGRA and scale uniforms with a uniform block * Setting the data again on program change is no longer needed * Optimize and resolve some warnings * Avoid redundant support buffer updates * Some optimizations to BindBuffers (now inlined) * Unify render scale arrays	2021-08-11 21:33:43 +02:00
gdkchan	d9d18439f6	Use a new approach for shader BRX targets (#2532 ) * Use a new approach for shader BRX targets * Make shader cache actually work * Improve the shader pattern matching a bit * Extend LDC search to predecessor blocks, catches more cases * Nit * Only save the amount of constant buffer data actually used. Avoids crashes on partially mapped buffers * Ignore Rd on predicate instructions, as they do not have a Rd register (catches more cases)	2021-08-11 20:59:42 +02:00
gdkchan	ff5df5d8a1	Support non-contiguous copies on I2M and DMA engines (#2473 ) * Support non-contiguous copies on I2M and DMA engines * Vector copy should start aligned on I2M * Nits * Zero extend the offset	2021-08-04 22:20:58 +02:00
riperiperi	4b60371e64	Return mapped buffer pointer directly for flush, WriteableRegion for textures (#2494 ) * Return mapped buffer pointer directly for flush, WriteableRegion for textures A few changes here to generally improve performance, even for platforms not using the persistent buffer flush. - Texture and buffer flush now return a ReadOnlySpan<byte>. It's guaranteed that this span is pinned in memory, but it will be overwritten on the next flush from that thread, so it is expected that the data is used before calling again. - As a result, persistent mappings no longer copy to a new array - rather the persistent map is returned directly as a Span<>. A similar host array is used for the glGet flushes instead of allocating new arrays each time. - Texture flushes now do their layout conversion into a WriteableRegion when the texture is not MultiRange, which allows the flush to happen directly into guest memory rather than into a temporary span, then copied over. This avoids another copy when doing layout conversion. Overall, this saves 1 data copy for buffer flush, 1 copy for linear textures with matching source/target stride, and 2 copies for block textures or linear textures with mismatching strides. * Fix tests * Fix array pointer for Mesa/Intel path * Address some feedback * Update method for getting array pointer.	2021-07-19 19:10:54 -03:00

1 2 3 4 5 ...

420 commits