summaryrefslogtreecommitdiffstats
path: root/libswscale/yuv2rgb_altivec.c
diff options
context:
space:
mode:
authordiego <diego@b3059339-0415-0410-9bf9-f77b7e298cf2>2008-07-04 13:49:45 +0000
committerdiego <diego@b3059339-0415-0410-9bf9-f77b7e298cf2>2008-07-04 13:49:45 +0000
commit1e95d9f4ef289418b8cf62303c2b1e8d53b6f880 (patch)
tree69d6c601d2a7ee13244000039abec4e4b4408d80 /libswscale/yuv2rgb_altivec.c
parent90db8eb5f2672b333e2cefd174b7f6c17b8d235f (diff)
downloadmpv-1e95d9f4ef289418b8cf62303c2b1e8d53b6f880.tar.bz2
mpv-1e95d9f4ef289418b8cf62303c2b1e8d53b6f880.tar.xz
spelling/grammar/wording overhaul
git-svn-id: svn://svn.mplayerhq.hu/mplayer/trunk@27190 b3059339-0415-0410-9bf9-f77b7e298cf2
Diffstat (limited to 'libswscale/yuv2rgb_altivec.c')
-rw-r--r--libswscale/yuv2rgb_altivec.c67
1 files changed, 36 insertions, 31 deletions
diff --git a/libswscale/yuv2rgb_altivec.c b/libswscale/yuv2rgb_altivec.c
index 3583e4bf65..d3b30e7ed5 100644
--- a/libswscale/yuv2rgb_altivec.c
+++ b/libswscale/yuv2rgb_altivec.c
@@ -21,63 +21,68 @@
*/
/*
-convert I420 YV12 to RGB in various formats,
- it rejects images that are not in 420 formats
- it rejects images that don't have widths of multiples of 16
- it rejects images that don't have heights of multiples of 2
-reject defers to C simulation codes.
+Convert I420 YV12 to RGB in various formats,
+ it rejects images that are not in 420 formats,
+ it rejects images that don't have widths of multiples of 16,
+ it rejects images that don't have heights of multiples of 2.
+Reject defers to C simulation code.
-lots of optimizations to be done here
+Lots of optimizations to be done here.
-1. need to fix saturation code, I just couldn't get it to fly with packs and adds.
- so we currently use max min to clip
+1. Need to fix saturation code. I just couldn't get it to fly with packs
+ and adds, so we currently use max/min to clip.
-2. the inefficient use of chroma loading needs a bit of brushing up
+2. The inefficient use of chroma loading needs a bit of brushing up.
-3. analysis of pipeline stalls needs to be done, use shark to identify pipeline stalls
+3. Analysis of pipeline stalls needs to be done. Use shark to identify
+ pipeline stalls.
MODIFIED to calculate coeffs from currently selected color space.
-MODIFIED core to be a macro which you spec the output format.
-ADDED UYVY conversion which is never called due to some thing in SWSCALE.
+MODIFIED core to be a macro where you specify the output format.
+ADDED UYVY conversion which is never called due to some thing in swscale.
CORRECTED algorithim selection to be strict on input formats.
-ADDED runtime detection of altivec.
+ADDED runtime detection of AltiVec.
ADDED altivec_yuv2packedX vertical scl + RGB converter
March 27,2004
PERFORMANCE ANALYSIS
-The C version use 25% of the processor or ~250Mips for D1 video rawvideo used as test
-The ALTIVEC version uses 10% of the processor or ~100Mips for D1 video same sequence
+The C version uses 25% of the processor or ~250Mips for D1 video rawvideo
+used as test.
+The AltiVec version uses 10% of the processor or ~100Mips for D1 video
+same sequence.
-720*480*30 ~10MPS
+720 * 480 * 30 ~10MPS
-so we have roughly 10clocks per pixel this is too high something has to be wrong.
+so we have roughly 10 clocks per pixel. This is too high, something has
+to be wrong.
-OPTIMIZED clip codes to utilize vec_max and vec_packs removing the need for vec_min.
+OPTIMIZED clip codes to utilize vec_max and vec_packs removing the
+need for vec_min.
-OPTIMIZED DST OUTPUT cache/dma controls. we are pretty much
-guaranteed to have the input video frame it was just decompressed so
-it probably resides in L1 caches. However we are creating the
-output video stream this needs to use the DSTST instruction to
-optimize for the cache. We couple this with the fact that we are
-not going to be visiting the input buffer again so we mark it Least
-Recently Used. This shaves 25% of the processor cycles off.
+OPTIMIZED DST OUTPUT cache/DMA controls. We are pretty much guaranteed to have
+the input video frame, it was just decompressed so it probably resides in L1
+caches. However, we are creating the output video stream. This needs to use the
+DSTST instruction to optimize for the cache. We couple this with the fact that
+we are not going to be visiting the input buffer again so we mark it Least
+Recently Used. This shaves 25% of the processor cycles off.
-Now MEMCPY is the largest mips consumer in the system, probably due
+Now memcpy is the largest mips consumer in the system, probably due
to the inefficient X11 stuff.
GL libraries seem to be very slow on this machine 1.33Ghz PB running
Jaguar, this is not the case for my 1Ghz PB. I thought it might be
-a versioning issues, however I have libGL.1.2.dylib for both
-machines. ((We need to figure this out now))
+a versioning issue, however I have libGL.1.2.dylib for both
+machines. (We need to figure this out now.)
-GL2 libraries work now with patch for RGB32
+GL2 libraries work now with patch for RGB32.
-NOTE quartz vo driver ARGB32_to_RGB24 consumes 30% of the processor
+NOTE: quartz vo driver ARGB32_to_RGB24 consumes 30% of the processor.
-Integrated luma prescaling adjustment for saturation/contrast/brightness adjustment.
+Integrated luma prescaling adjustment for saturation/contrast/brightness
+adjustment.
*/
#include <stdio.h>