FFmpeg

mirror of https://git.ffmpeg.org/ffmpeg.git synced 2024-09-21 05:46:55 +00:00

Author	SHA1	Message	Date
Hendrik Leppkes	2214207d04	Merge commit '8563f9887194b07c972c3475d6b51592d77f73f7' * commit '8563f9887194b07c972c3475d6b51592d77f73f7': x86: use emms after ff_int32_to_float_fmul_scalar_sse Merged-by: Hendrik Leppkes <h.leppkes@gmail.com>	2016-01-02 13:27:11 +01:00
Hendrik Leppkes	a9cd11b212	Merge commit 'f4f27e4cf1013c55b2c7df359ce8d58ee922662c' * commit 'f4f27e4cf1013c55b2c7df359ce8d58ee922662c': x86: zero extend the 32-bit length in int32_to_float_fmul_scalar implicitly Merged-by: Hendrik Leppkes <h.leppkes@gmail.com>	2016-01-02 13:23:25 +01:00
Hendrik Leppkes	d03da3e240	Merge commit '2008f76054906e9ff6bf744800af0e5a5bfe61be' * commit '2008f76054906e9ff6bf744800af0e5a5bfe61be': dca: remove unused decode_hf function and quant_d tables Merged-by: Hendrik Leppkes <h.leppkes@gmail.com>	2016-01-02 13:17:48 +01:00
Hendrik Leppkes	00e91d0676	Merge commit '5dfe4edad63971d669ae456b0bc40ef9364cca80' * commit '5dfe4edad63971d669ae456b0bc40ef9364cca80': x86_64: int32_to_float_fmul_scalar sign extend integer length Merged-by: Hendrik Leppkes <h.leppkes@gmail.com>	2016-01-02 10:46:18 +01:00
Janne Grunau	8563f98871	x86: use emms after ff_int32_to_float_fmul_scalar_sse Intel's Instruction Set Reference (as of September 2015) clearly states that cvtpi2ps switches to MMX state. Actual CPUs do not switch if the source is a memory location. The Instruction Set Reference from 1999 (Order Number 243191) describes this behaviour but all later versions I've seen have make no distinction whether MMX registers or memory is used as source. The documentation for the matching SSE2 instruction to convert to double (cvtpi2pd) was fixed (see the valgrind bug https://bugs.kde.org/show_bug.cgi?id=210264). It will take time to get a clarification and fixes in place. In the meantime it makes sense to change ff_int32_to_float_fmul_scalar_sse to be correct according to the documentation. The vast majority of users will have SSE2 so a change to the SSE version has little effect. Fixes fate-checkasm on x86 valgrind targets. Valgrind 'bug' reported as https://bugs.kde.org/show_bug.cgi?id=357059	2015-12-30 13:37:57 +01:00
Janne Grunau	f4f27e4cf1	x86: zero extend the 32-bit length in int32_to_float_fmul_scalar implicitly This reverts commit `5dfe4edad6`.	2015-12-29 11:42:51 +01:00
Alexandra Hájková	2008f76054	dca: remove unused decode_hf function and quant_d tables They were superseded with their integer equivalents. Rename integer decode_hf to decode_hf.	2015-12-24 13:58:18 +01:00
James Almer	d4c47333e1	x86/hevc_sao: add ff_hevc_sao_edge_filter_{8,16}_{10,12} Reviewed-by: Christophe Gisquet <christophe.gisquet@gmail.com> Signed-off-by: James Almer <jamrial@gmail.com>	2015-12-20 17:01:15 -03:00
James Almer	3ff2beff65	x86/hevc_sao: simplify sao_edge_filter 10/12bit Reviewed-by: Michael Niedermayer <michaelni@gmx.at> Reviewed-by: Christophe Gisquet <christophe.gisquet@gmail.com> Signed-off-by: James Almer <jamrial@gmail.com>	2015-12-20 16:45:37 -03:00
James Almer	34b2bd03cf	x86/hevc_sao: simplify sao_band_filter 10/12bit Reviewed-by: Michael Niedermayer <michaelni@gmx.at> Reviewed-by: Christophe Gisquet <christophe.gisquet@gmail.com> Signed-off-by: James Almer <jamrial@gmail.com>	2015-12-20 16:42:36 -03:00
Janne Grunau	5dfe4edad6	x86_64: int32_to_float_fmul_scalar sign extend integer length	2015-12-14 16:42:35 +01:00
Dave Yeo	b0b133b8c0	hevcdsp: use a macro for .rodata section fixes assembling on OS/2 Signed-off-by: Dave Yeo <dave.r.yeo@gmail.com> Signed-off-by: Luca Barbato <lu_zero@gentoo.org>	2015-12-11 16:19:30 +01:00
Kieran Kunhya	3f07f12f65	diracdec: Template DSP functions adding 10-bit versions	2015-12-10 18:25:02 +00:00
Anton Khirnov	e7078e842d	hevcdsp: add x86 SIMD for MC	2015-12-05 21:11:52 +01:00
Timothy Gu	4b80b895a9	pixblockdsp: x86: Condense diff_pixels_* to a shared macro Reviewed-by: Ronald S. Bultje <rsbultje@gmail.com> Reviewed-by: James Almer <jamrial@gmail.com>	2015-11-07 14:31:34 -08:00
Ganesh Ajjanagadde	38f4e973ef	all: fix -Wextra-semi reported on clang This fixes extra semicolons that clang 3.7 on GNU/Linux warns about. These were trigggered when built under -Wpedantic, which essentially checks for strict ISO compliance in numerous ways. Reviewed-by: Michael Niedermayer <michael@niedermayer.cc> Signed-off-by: Ganesh Ajjanagadde <gajjanagadde@gmail.com>	2015-10-24 17:58:17 -04:00
Ronald S. Bultje	52f84d82bd	videodsp: don't overread edges in vfix3 emu_edge. Fixes trac ticket 3226. Also see Andreas' analysis in https://bugs.debian.org/801745, which was very helpful.	2015-10-24 14:34:50 -04:00
Michael Niedermayer	ea5a1d1485	avcodec/x86/vc1dsp: Remove unused macro Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>	2015-10-22 21:13:42 +02:00
Carl Eugen Hoyos	775b84e30e	lavc/x86/vc1dsp_init: Fix compilation with --disable-yasm.	2015-10-22 11:37:42 +02:00
James Almer	73353af6e5	x86/Makefile: move decoder/encoder objects out of the subsystems section Signed-off-by: James Almer <jamrial@gmail.com>	2015-10-22 03:55:18 -03:00
Timothy Gu	ab5f43e634	vc1dsp: Port ff_vc1_put_ver_16b_shift2_mmx to yasm This function is only used within other inline asm functions, hence the HAVE_MMX_INLINE guard. Per recent discussions, we should not worry about the performance of inline asm-only builds.	2015-10-21 20:01:52 -07:00
Timothy Gu	98da061461	huffyuvencdsp: Cherry pick changes left out in the last commit Oops.	2015-10-21 12:42:33 -07:00
Timothy Gu	5e586e1bef	huffyuvencdsp: Add ff_diff_bytes_{sse2,avx2} SSE2 version 4%-35% faster than MMX depending on the width. AVX2 version 1%-13% faster than SSE2 depending on the width.	2015-10-21 12:25:32 -07:00
Timothy Gu	6b41b44149	huffyuvencdsp: Convert ff_diff_bytes_mmx to yasm Heavily based upon ff_add_bytes by Christophe Gisquet. Reviewed-by: James Almer <jamrial@gmail.com> Signed-off-by: Timothy Gu <timothygu99@gmail.com>	2015-10-20 18:24:54 -07:00
Timothy Gu	068e6cb732	huffyuvencdsp: Use intptr_t for width It is done this way in huffyuvdsp as well.	2015-10-19 16:57:33 -07:00
Timothy Gu	a079cbf458	x86: vc1dsp_mmx: Move yasm initiation steps to vc1dsp_init That's where all yasm initiation steps are. Also removes the overlap between the two files.	2015-10-19 16:52:52 -07:00
Timothy Gu	607f820ec7	x86: fpel: Remove erroneous ff_put_pixels8_mmxext prototype This function does not exist.	2015-10-19 16:52:37 -07:00
Timothy Gu	cb6f1f8bf9	x86: fpel: Move prototypes for 4-px block functions	2015-10-19 16:52:33 -07:00
James Almer	74a87ae210	x86/vp9itxfm: fix register clobbering in ff_vp9_idct_idct_4x4_add_12_sse2 Reviewed-by: Henrik Gramner <henrik@gramner.com> Signed-off-by: James Almer <jamrial@gmail.com>	2015-10-13 20:21:33 -03:00
Christophe Gisquet	74c414202f	x86: simple_idct10_template: use const This avoid going through constants.c while still sharing them with proresdsp.asm Reviewed-by: "Ronald S. Bultje" <rsbultje@gmail.com> Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>	2015-10-13 22:52:33 +02:00
Ronald S. Bultje	e578638382	vp9: use registers for constant loading where possible.	2015-10-13 11:06:01 -04:00
Ronald S. Bultje	408bb8556f	vp9: refactor itx coefficients and share between 8 and 10/12bpp.	2015-10-13 11:06:01 -04:00
Ronald S. Bultje	eb4b5ff738	vp9: add itxfm_add eob shortcuts to 10/12bpp functions. These aren't quite as helpful as the ones in 8bpp, since over there, we can use pmulhrsw, but here the coefficients have too many bits to be able to take advantage of pmulhrsw. However, we can still skip cols for which all coefs are 0, and instead just zero the input data for the row itx. This helps a few % on overall decoding speed.	2015-10-13 11:06:01 -04:00
Ronald S. Bultje	488fadebbc	vp9: add 10/12bpp idct_idct_32x32 sse2 SIMD version.	2015-10-13 11:06:00 -04:00
Ronald S. Bultje	3d0ca2fe89	vp9: 10/12bpp sse2 SIMD for iadst16.	2015-10-13 11:06:00 -04:00
Ronald S. Bultje	0e80265b0a	vp9: refactor 10/12bpp dc-only code in 4x4/8x8 and add to 16x16.	2015-10-13 11:06:00 -04:00
Ronald S. Bultje	1338fb79d4	vp9: add 10/12bpp sse2 SIMD version for idct_idct_16x16.	2015-10-13 11:06:00 -04:00
Ronald S. Bultje	cb054d061a	vp9: add 10/12bpp sse2 SIMD versions of iadst8x8.	2015-10-13 11:05:59 -04:00
Ronald S. Bultje	e0610787b2	vp9: add 10/12bpp sse2 SIMD for idct_idct_8x8.	2015-10-13 11:05:59 -04:00
Ronald S. Bultje	a35f6bdb38	vp9: add 12bpp sse2 versions of iadst4.	2015-10-13 11:05:59 -04:00
Ronald S. Bultje	235e76aeb8	vp9: initial attempt at a idct_idct_4x4 12bpp x86 simd (sse2) impl. The trouble with this function is that intermediates overflow 31+sign bits, so I've added some helpers (that will also be used in 10/12bpp 8x8, 16x16 and 32x32) to make that easier, basically emulating a half- assed pmaddqd using 2xpmaddwd. It's currently sse2-only, if anyone sees potential in adding ssse3, I'd love to hear it.	2015-10-13 11:05:58 -04:00
Ronald S. Bultje	f76423d097	vp9: add x86 simd (sse2/ssse3) for iadst4 10bpp functions.	2015-10-13 11:05:58 -04:00
Ronald S. Bultje	6b579cf547	vp9: add 10bpp simd (mmxext/ssse3) for idct_idct_4x4.	2015-10-13 11:05:58 -04:00
Ronald S. Bultje	1c3be32533	vp9: add 10/12bpp mmxext-optimized iwht_iwht_4x4 function.	2015-10-13 11:05:57 -04:00
Christophe Gisquet	b6594a9605	x86: dct-test: add more idcts In particular for 10 and 12 bits. Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>	2015-10-13 16:03:04 +02:00
Christophe Gisquet	7ece8b50b1	x86: simple_idct: 12bits versions On 12 frames of a 444p 12 bits DNxHR sequence, _put function: C: 78902 decicycles in idct, 262071 runs, 73 skips avx: 32478 decicycles in idct, 262045 runs, 99 skips Difference between the 2: stddev: 0.39 PSNR:104.47 MAXDIFF: 2 This is unavoidable and due to the scale factors used in the x86 version, which cannot match the C ones. In addition, the trick of adding an initial bias to the input of a pass can overflow, as the input coefficients are already 15bits, which is the maximum this function can handle. Overall, however, the omse on 12 bits samples goes from 0.16916 to 0.16883. Reducing rowshift by 1 improves to 0.0908, but causes overflows. Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>	2015-10-13 15:34:32 +02:00
Christophe Gisquet	4369b9dc7b	x86: simple_idct(_put): 10bits versions Modeled from the prores version. Clips to [0;1023] and is bitexact. Bitexactness requires to add offsets in different places compared to prores or C, and makes the function approximately 2% slower. For 16 frames of a DNxHD 4:2:2 10bits test sequence: C: 60861 decicycles in idct, 1048205 runs, 371 skips sse2: 27567 decicycles in idct, 1048216 runs, 360 skips avx: 26272 decicycles in idct, 1048171 runs, 405 skips The add version is not implemented, so the corresponding dsp function is set to NULL to make it clear in a code executing it. Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>	2015-10-13 13:32:21 +02:00
Christophe Gisquet	e652f69b35	x86: simple_idct10_template: fix overflow in pass When the input of a pass has 15 or 16 bits of precision (in particular the column pass), the addition of a bias to W4 may lead to overflows in the input to pmaddwd. This requires postponing the adding of the bias to after the first butterfly. To do so, the fact that m15, unused although zeroed, is exploited. In case the pass is safe, an address can be directly used, and the number of xmm regs can be decreased. Otherwise, the 32bits bias is loaded into it. Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>	2015-10-13 12:51:10 +02:00
Christophe Gisquet	e9a68b0316	x86: prores: templatize 10 bits simple_idct This should be reused for a generic simple_idct10 function. Requires a bit of trickery to declare common constants in C. Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>	2015-10-13 01:10:34 +02:00
James Almer	dab5f65b25	x86/takdsp: use arithmetic shift instructions p1 and p2 are int32_t. Reviewed-by: Ronald S. Bultje <rsbultje@gmail.com> Signed-off-by: James Almer <jamrial@gmail.com>	2015-10-09 23:52:39 -03:00

1 2 3 4 5 ...

2086 Commits