external_mesa3d.git - Unnamed repository; edit this file 'description' to name the repository.

	Commit message (Collapse)	Author	Age	Files	Lines
*	llvmpipe: enable z32s8x24 format	Roland Scheidegger	2013-05-18	1	-6/+0
\| \| \| \| \| \| \| \|	Now that we can handle it both for sampling and as depth/stencil enable it. Passes nearly all additional piglit tests which are now performed, with two exceptions (one being a framebuffer blit which fails for all other formats including stencil too as we don't support stencil blits, the other reporting a unexpected GL error so doesn't look to be llvmpipe's fault).
*	llvmpipe: handle z32s8x24 depth/stencil format	Roland Scheidegger	2013-05-18	7	-147/+252
\| \| \| \| \| \| \|	We need to split up the depth and stencil values in this case, and there's some new logic required to handle float depth and stencil simultaneously. Also make sure we get the 64bit zs clear values and masks propagated correctly.
*	llvmpipe: get rid of unused tiled/linear logic	Roland Scheidegger	2013-05-18	7	-713/+50
\| \| \| \| \| \| \| \| \|	We do rendering to linear color buffers for quite some time, and since switching to linear depth buffers all the tiled/linear logic was unused. So get rid of (most) of it - there's still some LAYOUT_NONE things and late allocation of resources which probably could be simplified. Reviewed-by: Jose Fonseca <jfonseca@vmware.com>
*	llvmpipe: fix bogus handling of first_layer when setting up texture sampling	Roland Scheidegger	2013-05-18	2	-14/+18
\| \| \| \| \| \| \| \| \| \| \| \|	The code avoided first_layer parameter in the sampler interface (and needing to do another calculation at runtime) by fixing up the base texture pointer instead. Unfortunately, this didn't actually work as we have mip-first texture layout so fixing up the base ptr by a fixed amount is very wrong if there are mipmaps present. The wrong offsets caused misrendering and crashes. Fix this by just adjusting the individual mip level offsets instead. Spotted by Jose. Reviewed-by: Jose Fonseca <jfonseca@vmware.com>
*	gallivm: Eliminate 8.8 fixed point intermediates from AoS sampling path.	José Fonseca	2013-05-17	1	-2/+2
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	This change was meant as a stepping stone to use PMADDUBSW SSSE3 instruction, but actually this refactoring by itself yields a 10% speedup on texture intensive shaders (e.g, Google Earth's ocean water w/o S3TC on a Ivy Bridge machine), while giving yielding exactly the same results, whereas PMADDUBSW only gave an extra 5%, at the expense of 2bits of precision in the interpolation. I belive that the speedup of this change comes from the reduced register pressure (as 8.8 fixed point intermediates take twice the space of 8bit unorm). Also, not dealing with 8.8 simplifies lp_bld_sample_aos.c code substantially -- it's no longer necessary to have code duplicated for low and high register halfs. Note about lp_build_sample_mipmap(): the path for num_quads > 1 is never executed (as it is faster on AVX to split the 256bit wide texture computation into two 128bit chunks, in order to leverage integer opcodes). This path might be useful in the future, so in order to verify this change did not break that path I had to apply this change: @@ -1662,11 +1662,11 @@ lp_build_sample_soa(struct gallivm_state gallivm, / * we only try 8-wide sampling with soa as it appears to * be a loss with aos with AVX (but it should work). * (It should be faster if we'd support avx2) / - if (num_quads == 1 \|\| !use_aos) { + if (/ num_quads == 1 \|\| ! / use_aos) { if (num_quads > 1) { if (mip_filter == PIPE_TEX_MIPFILTER_NONE) { LLVMValueRef index0 = lp_build_const_int32(gallivm, 0); / and then run texfilt mesademo: LP_NATIVE_VECTOR_WIDTH=256 ./texfilt Ran whole piglit without regressions. Reviewed-by: Roland Scheidegger <sroland@vmware.com>
*	radeon/llvm: Run standard optimization passes on conpute shader modules	Tom Stellard	2013-05-17	1	-0/+15
\| \| \| \| \| \|	The SROA and function inliner passes are espically important, because they optimize away unsupported features: functions and indirect private memory access.
*	r600g: fixup for MSAA texture support checking	Niels Ole Salscheider	2013-05-16	1	-1/+1
\| \| \| \|	Signed-off-by: Niels Ole Salscheider <niels_ole@salscheider-online.de>
*	llvmpipe: Temporary workaround to prevent segfault on array textures.	José Fonseca	2013-05-16	1	-0/+3
\|
*	ilo: emit 3DSTATE_STENCIL_BUFFER on GEN7+	Chia-I Wu	2013-05-16	2	-7/+21
\| \| \| \| \| \|	Whether HiZ is enalbed or not, separate stencil is supported and enforced on GEN7+. Now that we support separate stencil resources, we know how to emit 3DSTATE_STENCIL_BUFFER.
*	ilo: add support for stencil resources on GEN7+	Chia-I Wu	2013-05-16	8	-33/+545
\| \| \| \| \| \|	For allocations, we need to support stencil-only and separate stencil resources. For mapping, we need to support software tiling and packing/unpacking for separate stencil resources.
*	r600g: cleanup MSAA texture support checking	Marek Olšák	2013-05-15	7	-72/+21
\| \| \| \|	Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
*	r600g: rewrite FMASK allocation, fix FMASK texturing with 2 and 4 samples	Marek Olšák	2013-05-15	7	-36/+42
\| \| \| \| \| \| \| \| \| \| \| \|	This fixes and enables texturing with compressed MSAA colorbuffers on Evergreen and Cayman. For the first time, multisample textures work on Cayman. This requires the libdrm flag RADEON_SURF_FMASK. v2: require libdrm_radeon 2.4.45 Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
*	ilo: clean up transfer format conversion	Chia-I Wu	2013-05-15	1	-34/+48
\| \| \| \|	Map the bo directly, instead of calling transfer_map().
*	ilo: rework transfer mapping method choosing	Chia-I Wu	2013-05-15	1	-102/+133
\| \| \| \| \| \|	Always check if a bo is busy in choose_transfer_method() since we always need to map it in either map() or unmap(). Also determine how a bo is mapped in choose_transfer_method().
*	ilo: refactor transfer mapping	Chia-I Wu	2013-05-15	1	-27/+52
\| \| \| \| \|	Add tex_get_box_offset() to compute transfer offet from the pipe_box. Add tex_get_slice_stride() to compute slice stride for a transfer.
*	ilo: no writeback without PIPE_TRANSFER_WRITE	Chia-I Wu	2013-05-15	1	-0/+5
\| \| \| \|	We should not write staging data back when PIPE_TRANSFER_WRITE is not set.
*	ilo: minor cleanups for transfers	Chia-I Wu	2013-05-15	1	-41/+41
\| \| \| \|	Rename some functions and reorder some code.
*	ilo: simplify ilo_texture_get_slice_offset()	Chia-I Wu	2013-05-15	4	-55/+40
\| \| \| \|	Always return a tile-aligned offset. Also fix for W tiling.
*	draw: try to prevent overflows on index buffers	Zack Rusin	2013-05-14	6	-14/+27
\| \| \| \| \| \| \| \| \| \| \| \|	Pass in the size of the index buffer, when available, and use it to handle out of bounds conditions. The behavior in the case of an overflow needs to be the same as with other overflows in the vertex processing pipeline meaning that a vertex should still be generated but all attributes in it set to zero. Signed-off-by: Zack Rusin <zackr@vmware.com> Reviewed-by: José Fonseca <jfonseca@vmware.com> Reviewed-by: Roland Scheidegger <sroland@vmware.com>
*	draw: don't crash on vertex buffer overflow	Zack Rusin	2013-05-14	6	-11/+15
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	We would crash when stride was bigger than the size of the buffer. The correct behavior is to just fetch zero's in this case. Unfortunatly with user_buffer's there's no way to validate the size because currently we're just not getting it. Adjust the draw interface to pass the size along the mapped buffer, which works perfectly for buffer backed vertex_buffers and, in future, it will allow us to plumb user_buffer sizes through the same interface. Signed-off-by: Zack Rusin <zackr@vmware.com> Reviewed-by: José Fonseca <jfonseca@vmware.com> Reviewed-by: Roland Scheidegger <sroland@vmware.com>
*	radeonsi: update r600_get_llvm_processor_name for hainan	Alex Deucher	2013-05-14	1	-0/+1
\| \| \| \| \|	Signed-off-by: Alex Deucher <alexander.deucher@amd.com> Reviewed-by: Michel Dänzer <michel.daenzer@amd.com>
*	radeonsi: add support for hainan chips	Alex Deucher	2013-05-14	2	-0/+4
\| \| \| \| \| \| \|	Note: this is a candidate for the 9.1 branch Signed-off-by: Alex Deucher <alexander.deucher@amd.com> Reviewed-by: Michel Dänzer <michel.daenzer@amd.com>
*	r600g/sb: add missing cases for ARUBA chips	Vadim Girlin	2013-05-14	2	-0/+2
\| \| \| \|	Signed-off-by: Vadim Girlin <vadimgirlin@gmail.com>
*	r600g/sb: get rid of standard c++ streams	Vadim Girlin	2013-05-14	24	-545/+592
\| \| \| \| \| \| \| \| \| \| \| \|	Static initialization of internal libstdc++ data related to iostream causes segfaults with some apps. This patch replaces all uses of std::ostream and std::ostringstream in sb with custom lightweight classes. Prevents segfaults with ut2004demo and probably some other old apps. Signed-off-by: Vadim Girlin <vadimgirlin@gmail.com>
*	r600g/sb: separate bytecode decoding and parsing	Vadim Girlin	2013-05-14	6	-144/+163
\| \| \| \| \| \| \| \| \|	Parsing and ir construction is required for optimization only, it's unnecessary if we only need to print shader dump. This should make new disassembler more tolerant to any new features in the bytecode. Signed-off-by: Vadim Girlin <vadimgirlin@gmail.com>
*	ilo: rework ilo_texture	Chia-I Wu	2013-05-14	5	-766/+1027
\| \| \| \| \|	Use ilo_buffer for buffer resources and ilo_texture for texture resources. A major cleanup is necessitated by the separation.
*	ilo: rename ilo_resource to ilo_texture	Chia-I Wu	2013-05-14	7	-322/+322
\| \| \| \|	In preparation for the introduction of ilo_buffer.
*	ilo: move transfer-related functions to a new file	Chia-I Wu	2013-05-14	6	-450/+518
\| \| \| \| \| \|	Resource mapping is distinct from resource allocation, and is going to get more and more complex. Move the related functions to a new file to make the separation clear.
*	gallium: add PIPE_CAP_MAX_TEXTURE_BUFFER_SIZE for GL	Marek Olšák	2013-05-11	9	-0/+15
\| \| \| \| \| \|	v2: fix typo 65535 -> 65536 Reviewed-by: Brian Paul <brianp@vmware.com>
*	ilo: Initialize read_back in transfer_map_sys.	Vinson Lee	2013-05-10	1	-1/+1
\| \| \| \| \| \| \|	Fixes "Uninitialized scalar variable" defect reported by Coverity. Signed-off-by: Vinson Lee <vlee@freedesktop.org> Reviewed-by: Chia-I Wu <olvaffe@gmail.com>
*	r600g: increase array size for shader inputs and outputs	Marek Olšák	2013-05-10	2	-2/+4
\| \| \| \| \| \| \|	and add assertions to prevent buffer overflow. This fixes corruption of the r600_shader struct. NOTE: This is a candidate for the stable branches.
*	ilo: Add support for HW primitive restart.	Courtney Goeltzenleuchter	2013-05-10	3	-2/+194
\| \| \| \| \| \| \| \| \|	Now tells Gallium that ilo supports primitive restart. Updated ilo_draw_vbo to be able to check that the indexed primitive being rendered can actually be supported in HW. If not, will break up into individual prims similar to what Mesa does. [olv: a minor fix after rebasing and formatting]
*	svga: misc whitespace and comment fixes in svga_cmd.c	Brian Paul	2013-05-09	1	-82/+82
\|
*	ilo: add support for PIPE_FORMAT_ETC1_RGB8	Chia-I Wu	2013-05-09	2	-5/+61
\| \| \| \|	It is decompressed to and stored as PIPE_FORMAT_R8G8B8X8_UNORM on-the-fly.
*	ilo: support mapping with a staging system buffer	Chia-I Wu	2013-05-09	1	-0/+77
\| \| \| \| \|	It can be used for unpacking compressed texture on-the-fly or to support explicit transfer flushing.
*	ilo: allow for different mapping methods	Chia-I Wu	2013-05-09	1	-115/+187
\| \| \| \| \|	We want to or need to use a different mapping method when when the resource is busy, the bo format differs from the requested format, and etc.
*	ilo: allow bo format to differ from that requested	Chia-I Wu	2013-05-09	2	-14/+22
\| \| \| \| \|	For separate stencil buffer or formats not supported natively, the real format of the bo may differ from that requested.
*	i915: Use Y tiling for textures	Stéphane Marchesin	2013-05-08	1	-2/+7
\| \| \| \| \| \| \| \| \| \| \|	This basically reverts commit 2acc7193743199701f8f6d1877a59ece0ec4fa5b. With the previous change, we're not batchbuffer limited any longer. So we actually start seeing a performance difference between X and Y tiling. X tiling is funny because it is faster for screen-aligned quads but slower in games. So let's use Y tiling which is 10% faster overall.
*	i915g: Optimize batchbuffer sizes	Stéphane Marchesin	2013-05-08	1	-3/+5
\| \| \| \| \| \| \|	Now that we don't throttle at every batchbuffer, we can shrink the size of batchbuffers to achieve early flushing. This gives a significant speed boost in a lot of games (on the order of 20%).
*	i915g: Add more PIPE_CAP_* support	Stéphane Marchesin	2013-05-08	1	-0/+9
\|
*	ilo: remove our own type inference	Chia-I Wu	2013-05-08	1	-97/+27
\| \| \| \|	tgsi_opcode_infer_{src,dst}_type() works just fine.
*	ilo: use tgsi_util_get_texture_coord_dim()	Chia-I Wu	2013-05-08	3	-92/+4
\| \| \| \|	And remove toy_tgsi_get_texture_coord_dim().
*	nv50: initialize kick_notify callback in nv50_create	Bryan Cain	2013-05-07	1	-0/+1
\| \| \| \| \| \|	Fixes infinite loop on startup in Portal and Left 4 Dead 2. NOTE: This is a candidate for the 9.0 and 9.1 branches.
*	ilo: Add missing break statement in aos_tex TGSI_OPCODE_TEX2 case.	Vinson Lee	2013-05-07	1	-0/+1
\| \| \| \| \| \| \|	Fixes "Missing break in switch" defect reported by Coverity. Signed-off-by: Vinson Lee <vlee@freedesktop.org> Reviewed-by: Chia-I Wu <olvaffe@gmail.com>
*	r600g/sb: optimize some cases for CNDxx instructions	Vadim Girlin	2013-05-07	2	-5/+81
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	We can replace CNDxx with MOV (and possibly eliminate after propagation) in following cases: If src1 is equal to src2 in CNDxx instruction then the result doesn't depend on condition and we can replace the instruction with "MOV dst, src1". If src0 is const then we can evaluate the condition at compile time and also replace it with MOV. Signed-off-by: Vadim Girlin <vadimgirlin@gmail.com>
*	r600g/sb: fix memory leaks	Vadim Girlin	2013-05-07	2	-1/+7
\| \| \| \|	Signed-off-by: Vadim Girlin <vadimgirlin@gmail.com>
*	r600g/sb: fix kcache handling on r6xx	Vadim Girlin	2013-05-07	1	-1/+5
\| \| \| \| \| \| \| \| \|	Use the same limit for kcache constants in alu group on r6xx as on other chips (two const pairs). Relaxing this will require additional checks to make sure that all 4 consts in the group come from 2 kcache sets (clause limit), probably without noticeable improvements of shader performance. Signed-off-by: Vadim Girlin <vadimgirlin@gmail.com>
*	r600g/llvm: Parse config values in register / value pairs	Tom Stellard	2013-05-06	2	-4/+31
\| \| \| \|	Rather than relying on a predetermined order for the config values.
*	r600g/llvm: Don't feed LLVM output through r600_bytecode_build()	Tom Stellard	2013-05-06	4	-395/+21
\| \| \| \| \|	The LLVM backend emits raw ISA now, so we can just its output unmodified.
*	r600g/llvm: Don't emit CALL_FS for vertex shaders	Tom Stellard	2013-05-06	2	-8/+10
\| \| \| \|	The LLVM backend takes care of this now.