Ryujinx

Author	SHA1	Message	Date
gdkchan	11b437eafc	Fix DisplayInfo struct (#2708 )	2021-10-05 12:38:44 -03:00
riperiperi	fff48bb45a	Smaller initial size for ModifiedRangeList & directly inherit range list (#2663 ) This fixes a potential regression with the new range list changes, where the cost for creating new ones would be rather large due to creating a 1024 size array. Also reduces cost for range list inheritance by using the first existing range list as a base, rather than creating a new one then adding both lists to it. The growth size for the RangeList is now identical to its initial size. Every 32 elements was probably a little too common - now it is 1024 for most things and 8 for the buffer modified range list. The Unmapped and SyncMethod methods have been changed to ensure that they behave properly if the range list is set null. Cleaned up a few calls to use the null-conditional operator.	2021-10-04 15:38:59 -03:00
gdkchan	75f4b1ff2d	Relax sampler pool requirement (#2703 )	2021-10-04 14:35:28 -03:00
gdkchan	f7aaea4300	Unref frames before decoding with FFMPEG (#2704 )	2021-10-04 14:12:24 -03:00
riperiperi	d92fff541b	Replace CacheResourceWrite with more general "precise" write (#2684 ) * Replace CacheResourceWrite with more general "precise" write The goal of CacheResourceWrite was to notify GPU resources when they were modified directly, by looking up the modified address/size in a structure and calling a method on each resource. The downside of this is that each resource cache has to be queried individually, they all have to implement their own way to do this, and it can only signal to resources using the same PhysicalMemory instance. This PR adds the ability to signal a write as "precise" on the tracking, which signals a special handler (if present) which can be used to avoid unnecessary flush actions, or maybe even more. For buffers, precise writes specifically do not flush, and instead punch a hole in the modified range list to indicate that the data on GPU has been replaced. The downside is that precise actions must ignore the page protection bits and always signal - as they need to notify the target resource to ignore the sequence number optimization. I had to reintroduce the sequence number increment after I2M, as removing it was causing issues in rabbids kingdom battle. However - all resources modified by I2M are notified directly to lower their sequence number, so the problem is likely that another unrelated resource is not being properly updated. Thankfully, doing this does not affect performance in the games I tested. This should fix regressions from #2624. Test any games that were broken by that. (RF4, rabbids kingdom battle) I've also added a sequence number increment to ThreedClass.IncrementSyncpoint, as it seems to fix buffer corruption in OpenGL homebrew. (this was a regression from removing sequence number increment from constant buffer update - another unrelated resource thing) * Add tests. * Add XML docs for GpuRegionHandle * Skip UpdateProtection if only precise actions were called This allows precise actions to skip reprotection costs.	2021-09-29 02:27:03 +02:00
riperiperi	b6e093b0fc	Force copy when auto-deleting a texture with dependencies (#2687 ) When a texture is deleted by falling to the bottom of the AutoDeleteCache, its data is flushed to preserve any GPU writes that occurred. This ensures that the data appears in any textures recreated in the future, but didn't account for a texture that already existed with a copy dependency. This change forces copy dependencies to complete if a texture falls out from from the AutoDeleteCache. (not removed via overlap, as that would be wasted effort) Fixes broken lighting caused by pausing in SMO's Metro Kingdom. May fix some other issues.	2021-09-29 02:11:05 +02:00
gdkchan	fd7567a6b5	Only make render target 2D textures layered if needed (#2646 ) * Only make render target 2D textures layered if needed * Shader cache version bump * Ensure topology is updated on channel swap	2021-09-29 01:55:12 +02:00
FICTURE7	312be74861	Optimize `HybridAllocator` (#2637 ) * Store constant `Operand`s in the `LocalInfo` Since the spill slot and register assigned is fixed, we can just store the `Operand` reference in the `LocalInfo` struct. This allows skipping hitting the intern-table for a look up. * Skip `Uses`/`Assignments` management Since the `HybridAllocator` is the last pass and we do not care about uses/assignments we can skip managing that when setting destinations or sources. * Make `GetLocalInfo` inlineable Also fix a possible issue where with numbered locals. See or-assignment operator in `SetVisited(local)` before patch. * Do not run `BlockPlacement` in LCQ With the host mapped memory manager, there is a lot less cold code to split from hot code. So disabling this in LCQ gives some extra throughput - where we need it. * Address Mou-Ikkai's feedback * Apply suggestions from code review Co-authored-by: VocalFan <45863583+Mou-Ikkai@users.noreply.github.com> * Move check to an assert Co-authored-by: VocalFan <45863583+Mou-Ikkai@users.noreply.github.com>	2021-09-29 01:38:37 +02:00
riperiperi	1ae690ba2f	Use normal memory store path for DC ZVA (#2693 ) Seems like this is used as an optimized way to clear memory in homebrew applications. Unfortunately, calling the software fallback method every 8 bytes was not very optimal. The existing EmitStore is used by passing in ZR as the register to get a 0 write.	2021-09-29 01:21:30 +02:00
Ac_K	33dc4c9ce4	clkrst: Stub/Implement IClkrstManager and IClkrstSession calls (#2692 ) This PR stubs and implements some clkrst call because they are used to overclock the Switch hardware and it's pointless in our case as we emulate the system. Everything was done checked by RE. Fixes #2686	2021-09-29 01:03:35 +02:00
gdkchan	f4f496cb48	NVDEC (H264): Use separate contexts per channel and decode frames in DTS order (#2671 ) * Use separate NVDEC contexts per channel (for FFMPEG) * Remove NVDEC -> VIC frame override hack * Add missing bottom_field_pic_order_in_frame_present_flag * Make FFMPEG logging static * nit: Remove empty lines * New FFMPEG decoding approach -- call h264_decode_frame directly, trim surface cache to reduce memory usage * Fix case * Silence warnings * PR feedback * Per-decoder rather than per-codec ownership of surfaces on the cache	2021-09-29 00:43:40 +02:00
FICTURE7	0d23504e30	Fix PTC count table relocation patching (#2666 ) Fix an issue introduced in #2190 where by 2 different count table entry addresses were used for LCQ functions. E.g: ```asm .L1: mov rbp,COUNT_TABLE_0 ;; This gets an address. mov ebp,[rbp] lea esi,[rbp+1] mov rdi,COUNT_TABLE_1 ;; This gets another address. mov [rdi],esi cmp ebp,64h je near .L34 ``` This caused LCQ functions to not tier up when they're loaded from the PTC cache. This does not happen when they're freshly compiled. This PR fixes the issue by ensuring only a single counter is created per translation.	2021-09-29 00:28:34 +02:00
Ac_K	79c854dd2e	irs: Stub some service calls (#2665 ) This PR stubs some irs service calls which are needed to get some games playable or at least bootable since we don't support IR data throught real JoyCon for now. - Stubs `IIrSensorServer` `StopImageProcessor`, `RunMomentProcessor`, `RunClusteringProcessor`, `RunImageTransferProcessor`, `GetImageTransferProcessorState`, `RunTeraPluginProcessor`. All calls are a bit checked by RE. Closes #2267, #2248, #2126 Night Vision and SpyAlarm are now bootable (but still unplayable due to the lack of the IR data):	2021-09-29 00:10:10 +02:00
gdkchan	83bdafccda	Share scales array for graphics and compute (#2653 )	2021-09-28 23:52:27 +02:00
VocalFan	405840a24b	Quick README update for game compatibility. (#2694 )	2021-09-28 23:26:45 +02:00
riperiperi	7c5ead1c19	Fast path for Inline2Memory buffer write that skips write tracking (#2624 ) * Fast path for Inline2Memory buffer write This PR adds a method to PhysicalMemory that attempts to write all cached resources directly, so that memory tracking can be avoided. The goal of this is both to avoid flushing buffer data, and to avoid raising the sequence number when data is written, which causes buffer and texture handles to be re-checked. This currently only targets buffers, with a side check on textures that falls back to a tracked write if any exist within the target range. It's not expected to write textures from here - this is just a mechanism to protect us if someone does decide to do that. It's possible to add a fast path for this in future (and for ShaderCache, once that starts using tracking) The forced read before inline2memory begins has been skipped, as the data is fully written when the transfer is completed anyways. This allows us to flush on read in emergency situations, but still write the new data over the flushed data. Improves performance on Xenoblade 2 and DE, which was flushing buffer data on the GPU thread when trying to write compute data. May improve performance in other games that write SSBOs from compute, and update data in the same/nearby pages often. Super Smash Bros Ultimate should probably be tested to make sure the vertex explosions haven't returned, as I think that's what this AdvanceSequence was for. * ForceDirty before write, to make sure data does not flush over the new write	2021-09-19 15:09:53 +02:00
riperiperi	db97b1d7d2	Implement and use an Interval Tree for the MultiRangeList (#2641 ) * Implement and use an Interval Tree for the MultiRangeList * Feedback * Address Feedback * Missed this somehow	2021-09-19 14:55:07 +02:00
gdkchan	f08a280ade	Use shader subgroup extensions if shader ballot is not supported (#2627 ) * Use shader subgroup extensions if shader ballot is not supported * Shader cache version bump + cleanup * The type is still required on the table	2021-09-19 14:38:39 +02:00
riperiperi	7379bc2f39	Array based RangeList that caches Address/EndAddress (#2642 ) * Array based RangeList that caches Address/EndAddress In isolation, this was more than 2x faster than the RangeList that checks using the interface. In practice I'm seeing much better results than I expected. The array is used because checking it is slightly faster than using a list, which loses time to struct copies, but I still want that data locality. A method has been added to the list to update the cached end address, as some users of the RangeList currently modify it dynamically. Greatly improves performance in Super Mario Odyssey, Xenoblade and any other GPU limited games. * Address Feedback	2021-09-19 14:22:26 +02:00
riperiperi	b0af010247	Set texture/image bindings in place rather than allocating and passing an array (#2647 ) * Remove allocations for texture bindings and state * Rent rather than stackalloc + copy A bit faster.	2021-09-19 14:03:05 +02:00
Mary	32c09af71a	amadeus: Fix regression from #2654 on ListAudioDeviceName	2021-09-19 13:42:16 +02:00
Ac_K	40d1acd198	vi: Unify resolutions values and accurate implementation of them. (#2640 ) * vi: Unify resolutions values and accurate implementation of them. To continue what was made in #2618, I've REd `vi` service a bit. Now values and checks related to displays are more accurate. - `am` GetDefaultDisplayResolution / GetDefaultDisplayResolutionChangeEvent have more informations on what the service does. - `vi:u/vi:m/vi:s` GetDisplayService are now accurate. - `IApplicationDisplay` GetRelayService, GetSystemDisplayService, GetManagerDisplayService, GetIndirectDisplayTransactionService, ListDisplays, OpenDisplay, OpenDefaultDisplay, CloseDisplay, GetDisplayResolution are now properly implemented. - Some other calls are cleaned or have extra checks accordingly to RE. Additionnaly, `IFriendService` have some wrong aligned things, and `pm:info` service placeholder was missing. * just use _openedDisplayInfo.Remove() * use context.Memory.Fill() * fix some casting * remove unneeded comment * cleanup * uses TryAdd * displayId > ulong * GetDisplayResolution > ulong * UL	2021-09-19 12:57:39 +02:00
Mary	e17eb7bfaf	amadeus: Update to REV10 (#2654 ) * amadeus: Update to REV10 This implements all the changes made with REV10 on 13.0.0. * Address Ack's comment * Address gdkchan's comment	2021-09-19 12:29:19 +02:00
mpnico	fe9d5a1981	Fix problems added by Pause (#2645 ) * Disable Pause/Resume menu instead of trying to hide them * Fix Resume menu being active before renderer starts * Fix emulator not being able to close properly	2021-09-18 14:31:44 +02:00
Ac_K	d327e809c9	gui: Hotfix for FileChooserNative during section extraction (#2644 ) Fix a regression introduced in #2633, FileChooserNative parent can't be set to null because it's running in modal.	2021-09-16 00:09:48 +02:00
MutantAura	843401635a	Adjustments to framerate metric and addition of frametime (#2638 ) * Adjust framerate data and add frametime * Update PerformanceStatistics.cs * Revert deletions of average framerate * Update Ryujinx.csproj * Remove separate GTK column * Increase FPS precision * general cleanup * even generaler cleanup * fix dumb * Remove legacy code * Update PerformanceStatistics.cs * Update PerformanceStatistics.cs	2021-09-15 02:26:10 +02:00
Michael Gielda	fb2e61a435	Add Linux Unicorn patch + desc. (#2609 )	2021-09-15 01:47:10 +02:00
Ac_K	5d08e9b495	hos: Cleanup the project (#2634 ) * hos: Cleanup the project Since a lot of changes has been done on the HOS project, there are some leftover here and there, or class just used in one service, things at wrong places, and more. This PR fixes that, additionnally to that, I've realigned some vars because I though it make the code more readable. * Address gdkchan feedback * addresses Thog feedback * Revert ElfSymbol	2021-09-15 01:24:49 +02:00
Ac_K	3f2486342b	gui: Replace FileChooserDialog by FileChooserNative (#2633 ) We currently use the FileChooser from GTK, which is a bit mess. Instead of it we could use the native FileChooser from all specifics OS. This is what this PR attempt to fix. It could be nice to get a test under linux since I've only tested it under Windows without any issues. Fixes #2584	2021-09-14 23:52:08 +02:00
FICTURE7	a9343c9364	Refactor `PtcInfo` (#2625 ) * Refactor `PtcInfo` This change reduces the coupling of `PtcInfo` by moving relocation tracking to the backend. `RelocEntry`s remains as `RelocEntry`s through out the pipeline until it actually needs to be written to the PTC streams. Keeping this representation makes inspecting and manipulating relocations after compilations less painful. This is something I needed to do to patch relocations to 0 to diff dumps. Contributes to #1125. * Turn `Symbol` & `RelocInfo` into readonly structs * Add documentation to `CompiledFunction` * Remove `Compiler.Compile<T>` Remove `Compiler.Compile<T>` and replace it by `Map<T>` of the `CompiledFunction` returned.	2021-09-14 01:23:37 +02:00
gdkchan	ac4ec1a015	Account for negative strides on DMA copy (#2623 ) * Account for negative strides on DMA copy * Should account for non-zero Y	2021-09-11 22:54:18 +02:00
gdkchan	016fc64b3d	Implement GetVaRegions on nvservices (#2621 ) * Implement GetVaRegions on nvservices * This would just result in 0	2021-09-11 22:39:02 +02:00
gdkchan	a4089fc878	Report 1080p resolution when in docked mode (#2618 )	2021-09-11 22:24:10 +02:00
mpnico	117e32a6ff	Implement a "Pause Emulation" option & hotkey (#2428 ) * Add a "Pause Emulation" option and hotkey Closes Ryujinx#1604 * Refactoring how pause is handled * Applied suggested changes from review * Applied suggested fixes * Pass correct suspend type to threads for suspend/resume * Fix NRE after stoping emulation * Removing SimulateWakeUpMessage call after resuming emulation * Skip suspending non game process * Pause the tickCounter in the ExecutionContext * Refactoring tickCounter pause/resume as suggested * Fix Config migration to add pause hotkey * Fixed pausing only application threads * Fix exiting emulator while paused * Avoid pause/resume while already paused/resumed * Cleanup unused code * Avoid restarting audio if stopping emulation while in pause. * Added suggested changes * Fix ConfigurationState	2021-09-11 22:08:25 +02:00
riperiperi	b0e410a828	Lift textures in the AutoDeleteCache for all modifications. (#2615 ) * Lift textures in the AutoDeleteCache for all modifications. Before, this would only apply to render targets and texture blit. Now it applies to image stores, the fast dma copy path and any other type of modification. Image store always at least has one reference in the texture pool, so the function of the AutoDeleteCache keeping textures _alive_ is not useful, but a very important function for a while has been its use to flush textures in order of modification when they are dereferenced, so that their data is not lost. Before, textures populated using image stores were being dereferenced and reloaded as garbage. Now, when these textures are dereferenced, their data will be put back into memory, and everything stays intact. Fixes lighting breaking when switching levels in THPS1+2, and potentially some more UE4 games. I've tested a bunch more games for regressions and performance impact, but they all seem fine. * Lift copy srcTexture so that it doesn't remain referenceless * Perform lift before reference count change on unbind. It's important to lift on unbind as that is the moment the texture was truly last modified, but definitely not after releasing every single reference.	2021-09-11 21:52:54 +02:00
Agustin Insua	197f587802	Fix GTK3 mapping for single quote key (#2612 )	2021-09-11 21:32:36 +02:00
Agustin Insua	bcbe6ef6cd	Update game metadata when stopping emulation (#2610 ) * Update game metadata when stopping emulation * Fix formatting	2021-09-11 21:16:48 +02:00
bobhope	830d1f097d	Remove file error popup (#2547 ) * Added check to detect if application file is more than zero bytes long * Removed file error popup * Removed unnecessary usings * Added empty lines	2021-09-11 20:59:11 +02:00
riperiperi	f0b00c1ae9	Fix TXQ for 3D textures. (#2613 ) * Fix TXQ for 3D textures. Assumes the texture is 3D if the component mask contains Z. This fixes a bug in UE4 games where parts of the map had garbage pointers to lighting voxels, as the lookup 3D texture was not being initialized. Most notable game is THPS1+2. May need another PR to keep image store data alive and properly flush it in order using the AutoDeleteCache. * Get sampler type for TextureSize from bound textures.	2021-09-02 00:17:43 -03:00
riperiperi	142cededd4	Implement Shader Instructions SUATOM and SURED (#2090 ) * Initial Implementation * Further improvements (no support for float/64-bit types) * Merge atomic and reduce instructions, add missing format switch * Fix rebase issues. * Not used. * Whoops. Fixed. * Partial implementation of inc/dec, cleanup and TODOs * Remove testing path * Address Feedback	2021-08-31 02:51:57 -03:00
gdkchan	416dc8fde4	Fix out-of-bounds shader thread shuffle (#2605 ) * Fix out-of-bounds shader thread shuffle * Shader cache version bump	2021-08-30 14:02:40 -03:00
gdkchan	82cefc8dd3	Handle indirect draw counts with non-zero draw starts properly (#2593 )	2021-08-29 16:52:38 -03:00
riperiperi	15e7fe3ac9	Avoid deleting textures when their data does not overlap. (#2601 ) * Avoid deleting textures when their data does not overlap. It's possible that while two textures start and end addresses indicate an overlap, that the actual data contained within them is sparse due to a layer stride. One such possibility is array slices of a cubemap at different mip levels - they overlap on a whole, but the actual texture data fills the gaps between each other's layers rather than actually overlapping. This fixes issues with UE4 games having incorrect lighting (solid white screen or really dark shadows). There are still remaining issues with games that use the 3D texture prebaked lighting, such as THPS1+2. This PR also fixes a bug with TexturePool's resized texture handling where the base level in the descriptor was not considered. * AllRegions granularity for 3d textures is now by level rather than by slice. * Address feedback	2021-08-29 16:22:13 -03:00
riperiperi	54adc5f9fb	Ensure that all threads wait for a read tracking action to complete. (#2597 ) * Lock around tracking action consume + execute. Not particularly fast. * Lock around preaction registration and use * Create a lock object * Nit	2021-08-29 16:03:41 -03:00
riperiperi	76e8f9ac87	Only reupload the texture scale array if it changes. (#2595 ) * Only reupload the texture scale array if it changes. Before, this would be called all the time if any shader needed a scale value. The cost of doing this has increased with threaded-gal, as the scale array is copied to a span pool, and it's was called on pretty much every draw sometimes. This improves GPU performance in games, scaled or not. Most affected game seems to be Xenoblade Chronicles: Definitive Edition. * Just use = instead of \|=	2021-08-27 17:08:30 -03:00
gdkchan	ee1038e542	Initial support for shader attribute indexing (#2546 ) * Initial support for shader attribute indexing * Support output indexing too, other improvements * Fix order * Address feedback	2021-08-27 01:44:47 +02:00
riperiperi	ec3e848d79	Add a Multithreading layer for the GAL, multi-thread shader compilation at runtime (#2501 ) * Initial Implementation About as fast as nvidia GL multithreading, can be improved with faster command queuing. * Struct based command list Speeds up a bit. Still a lot of time lost to resource copy. * Do shader init while the render thread is active. * Introduce circular span pool V1 Ideally should be able to use structs instead of references for storing these spans on commands. Will try that next. * Refactor SpanRef some more Use a struct to represent SpanRef, rather than a reference. * Flush buffers on background thread * Use a span for UpdateRenderScale. Much faster than copying the array. * Calculate command size using reflection * WIP parallel shaders * Some minor optimisation * Only 2 max refs per command now. The command with 3 refs is gone. 😌 * Don't cast on the GPU side * Remove redundant casts, force sync on window present * Fix Shader Cache * Fix host shader save. * Fixup to work with new renderer stuff * Make command Run static, use array of delegates as lookup Profile says this takes less time than the previous way. * Bring up to date * Add settings toggle. Fix Muiltithreading Off mode. * Fix warning. * Release tracking lock for flushes * Fix Conditional Render fast path with threaded gal * Make handle iteration safe when releasing the lock This is mostly temporary. * Attempt to set backend threading on driver Only really works on nvidia before launching a game. * Fix race condition with BufferModifiedRangeList, exceptions in tracking actions * Update buffer set commands * Some cleanup * Only use stutter workaround when using opengl renderer non-threaded * Add host-conditional reservation of counter events There has always been the possibility that conditional rendering could use a query object just as it is disposed by the counter queue. This change makes it so that when the host decides to use host conditional rendering, the query object is reserved so that it cannot be deleted. Counter events can optionally start reserved, as the threaded implementation can reserve them before the backend creates them, and there would otherwise be a short amount of time where the counter queue could dispose the event before a call to reserve it could be made. * Address Feedback * Make counter flush tracked again. Hopefully does not cause any issues this time. * Wait for FlushTo on the main queue thread. Currently assumes only one thread will want to FlushTo (in this case, the GPU thread) * Add SDL2 headless integration * Add HLE macro commands. Co-authored-by: Mary <mary@mary.zone>	2021-08-27 00:31:29 +02:00
Mary	501c3d5cea	Implement MSR instruction for A32 (#2585 ) * Implement MSR instruction Fix #1342. Now Pocket Rumble is playable. * Address gdkchan's comments * Address gdkchan's comments * Address gdkchan's comment	2021-08-27 00:07:44 +02:00
mpnico	8e1adb95cf	Add support for HLE macros and accelerate MultiDrawElementsIndirectCount #2 (#2557 ) * Add support for HLE macros and accelerate MultiDrawElementsIndirectCount * Add missing barrier * Fix index buffer count * Add support check for each macro hle before use * Add missing xml doc Co-authored-by: gdkchan <gab.dark.100@gmail.com>	2021-08-26 23:50:28 +02:00
VocalFan	5cab8ea4ad	Fix Unicorn Warnings (#2575 )	2021-08-26 23:34:24 +02:00

... 4 5 6 7 8 ...

2096 commits