百翔网络博客

含有HTML标签格式的文章简介截取功能实现

[ 2009年7月7日 ]

   由于上周在blog项目中要实现形如sina格式的文章摘要功能（就是在所见即所得HTML代码编辑器上编辑文章内容，在文章摘要中根据长度要求截取一部分的文章内容。截取后的内容还要含有HTML而且摘要的格式要和原文中没被截取的内容格式一样）。就baidu查找相关的功能代码可是找了一天都没有找到比较适合我要求的代码。然而就在今天快下班的前2个小时我终于找到了我想要的高能代码。
   看到网上还有很多朋友在寻找这样的功能代码。今天特意把我找到的代码粘贴出来，希望能给还找这样功能代码的朋友一点帮助。以下是原代码：
<?
/**
* Truncates text.
*
* Cuts a string to the length of $length and replaces the last characters
* with the ending if the text is longer than length.
*
* @param string $text String to truncate.
* @param integer $length Length of returned string, including ellipsis.
* @param string $ending Ending to be appended to the trimmed string.
* @param boolean $exact If false, $text will not be cut mid-word
* @param boolean $considerHtml If true, HTML tags would be handled correctly
* @return string Trimmed string.
*/
    function truncate($text, $length = 100, $ending = '...', $exact = true, $considerHtml = false) {
        if ($considerHtml) {
            // if the plain text is shorter than the maximum length, return the whole text
            if (strlen(preg_replace('/<.*?>/', '', $text)) <= $length) {
                return $text;
            }
            // splits all html-tags to scanable lines
            preg_match_all('/(<.+?>)?([^<>]*)/s', $text, $lines, PREG_SET_ORDER);
            $total_length = strlen($ending);
            $open_tags = array();
            $truncate = '';
            foreach ($lines as $line_matchings) {
                // if there is any html-tag in this line, handle it and add it (uncounted) to the output
                if (!empty($line_matchings[1])) {
                    // if it's an "empty element" with or without xhtml-conform closing slash (f.e. <br/>)
                    if (preg_match('/^<(\s*.+?\/\s*|\s*(img|br|input|hr|area|base|basefont|col|frame|isindex|link|meta|param)(\s.+?)?)>$/is', $line_matchings[1])) {
                        // do nothing
                    // if tag is a closing tag (f.e. </b>)
                    } else if (preg_match('/^<\s*\/([^\s]+?)\s*>$/s', $line_matchings[1], $tag_matchings)) {
                        // delete tag from $open_tags list
                        $pos = array_search($tag_matchings[1], $open_tags);
                        if ($pos !== false) {
                            unset($open_tags[$pos]);
                        }
                    // if tag is an opening tag (f.e. <b>)
                    } else if (preg_match('/^<\s*([^\s>!]+).*?>$/s', $line_matchings[1], $tag_matchings)) {
                        // add tag to the beginning of $open_tags list
                        array_unshift($open_tags, strtolower($tag_matchings[1]));
                    }
                    // add html-tag to $truncate'd text
                    $truncate .= $line_matchings[1];
                }
                // calculate the length of the plain text part of the line; handle entities as one character
                $content_length = strlen(preg_replace('/&[0-9a-z]{2,8};|&#[0-9]{1,7};|[0-9a-f]{1,6};/i', ' ', $line_matchings[2]));
                if ($total_length+$content_length> $length) {
                    // the number of characters which are left
                    $left = $length - $total_length;
                    $entities_length = 0;
                    // search for html entities
                    if (preg_match_all('/&[0-9a-z]{2,8};|&#[0-9]{1,7};|[0-9a-f]{1,6};/i', $line_matchings[2], $entities, PREG_OFFSET_CAPTURE)) {
                        // calculate the real length of all entities in the legal range
                        foreach ($entities[0] as $entity) {
                            if ($entity[1]+1-$entities_length <= $left) {
                                $left--;
                                $entities_length += strlen($entity[0]);
                            } else {
                                // no more characters left
                                break;
                            }
                        }
                    }
                    $truncate .= substr($line_matchings[2], 0, $left+$entities_length);
                    // maximum lenght is reached, so get off the loop
                    break;
                } else {
                    $truncate .= $line_matchings[2];
                    $total_length += $content_length;
                }
                // if the maximum length is reached, get off the loop
                if($total_length>= $length) {
                    break;
                }
            }
        } else {
            if (strlen($text) <= $length) {
                return $text;
            } else {
                $truncate = substr($text, 0, $length - strlen($ending));
            }
        }
        // if the words shouldn't be cut in the middle...
        if (!$exact) {
            // ...search the last occurance of a space...
            $spacepos = strrpos($truncate, ' ');
            if (isset($spacepos)) {
                // ...and cut the text in this position
                $truncate = substr($truncate, 0, $spacepos);
            }
        }
        // add the defined ending to the text
        $truncate .= $ending;
        if($considerHtml) {
            // close all unclosed html-tags
            foreach ($open_tags as $tag) {
                $truncate .= '</' . $tag . '>';
            }
        }
        return $truncate;
    }
?>

代码的原文网址是：http://www.gsdesign.ro/blog/cut-html-string-without-breaking-the-tags/#comment-7211

asp.net开源CMS汇总 (2009-6-25 13:10:10)

该学Java或.NET？ (2009-6-18 14:55:27)

flash制作工具大全 (2009-5-29 2:18:6)

4个电脑局域网一个共享的主机登陆失败:用户帐户限制,可能的原因可能包括不允许空密码 (2009-5-15 14:41:33)

PHP安全之错误报告 (2009-4-7 0:47:47)

PHP经验与心得 (2009-4-7 0:41:29)

几个常用CMS模板下载 (2009-2-12 23:15:55)

新云网站内容管理系统 v4.0.0.1230下载(2008.12.30) (2009-2-12 23:14:18)

申请google adsense帐户注意事项 (2009-2-9 15:55:39)

马云再抛“大招聘”今年招5000人 (2009-2-9 10:51:16)

点击这里获取该日志的TrackBack引用地址

发布:xunzhaohaizi | 分类:网站编程 | 评论:0 | 引用:0 | 浏览:

含有HTML标签格式的文章简介截取功能实现

Calendar

赞助商广告

最新评论及回复

最近发表