且构网

分享程序员开发的那些事...
且构网 - 分享程序员编程开发的那些事

将SQL_Latin1_General_CP1_CI_AS编码为UTF-8

更新时间:2023-11-27 21:47:40

我发现如何解决它,所以希望这将有助于某人。



首先,SQL_Latin1_General_CP1_CI_AS是CP-1252和UTF-8的奇怪组合。
基本的字符是CP-1252,所以这就是为什么我只需要做的就是UTF-8,一切正常。亚洲和其他UTF-8字符是以2个字节编码的,php pdo_mssql驱动程序似乎讨厌不同长度的字符,所以似乎做了一个CAST到varchar(而不是nvarchar),然后所有的2个字节的字符都成为问号(' ?')。



我通过将其转换为二进制来修复它,然后用php重建文本:

  SELECT CAST(MY_COLUMN AS VARBINARY(MAX))FROM MY_TABLE; 

在php:

  //二进制到十六进制
$ hex = bin2hex($ bin);

//然后从十六进制到字符串
$ str =;
($ i = 0; $ i< strlen($ hex)-1; $ i + = 2)
{
$ str。= chr(hexdec($ hex [$ i] 。$六角[$ I + 1]));
}
//然后从UCS-2LE / SQL_Latin1_General_CP1_CI_AS(这是DB中的列格式)到UTF-8
$ str = iconv('UCS-2LE','UTF- 8',$ str);


I'm generating a XML file with PHP using DomDocument and I need to handle asian characters. I'm pulling data from the MSSQL2008 server using the pdo_mssql driver and I apply utf8_encode() on the XML attribute values. Everything works fine as long as there's no special characters.

The server is MS SQL Server 2008 SP3

The database, table and column collation are all SQL_Latin1_General_CP1_CI_AS

I'm using PHP 5.2.17

Here's my PDO object:

$pdo = new PDO("mssql:host=MyServer,1433;dbname=MyDatabase", user123, password123);

My query is a basic SELECT.

I know storing special characters into SQL_Latin1_General_CP1_CI_AS columns isn't great, but ideally it would be nice to make it work without changing it, because other non-PHP programs already use that column and it works fine. In SQL Server Management Studio I can see the asian characters correctly.

Considering all the details above, how should I process the data?

I found how to solve it, so hopefully this will be helpful to someone.

First, SQL_Latin1_General_CP1_CI_AS is a strange mix of CP-1252 and UTF-8. The basic characters are CP-1252, so this is why all I had to do was UTF-8 and everything worked. The asian and other UTF-8 characters are encoded on 2 bytes and the php pdo_mssql driver seems to hate varying length characters so it seems to do a CAST to varchar (instead of nvarchar) and then all the 2 byte characters become question marks ('?').

I fixed it by casting it to binary and then I rebuild the text with php:

SELECT CAST(MY_COLUMN AS VARBINARY(MAX)) FROM MY_TABLE;

In php:

//Binary to hexadecimal
$hex = bin2hex($bin);

//And then from hex to string
$str = "";
for ($i=0;$i<strlen($hex) -1;$i+=2)
{
    $str .= chr(hexdec($hex[$i].$hex[$i+1]));
}
//And then from UCS-2LE/SQL_Latin1_General_CP1_CI_AS (that's the column format in the DB) to UTF-8
$str = iconv('UCS-2LE', 'UTF-8', $str);