From: "Nellis, Kenneth" <Kenneth.Nellis@acs-inc.com>
To: <cygwin-talk@cygwin.com>
Subject: internationalized text processing
Date: Mon, 23 May 2011 13:44:00 -0000 [thread overview]
Message-ID: <2BF01EB27B56CC478AD6E5A0A28931F202910C7C@A1DAL1SWPES19MB.ams.acs-inc.net> (raw)
Never posted here, but somewhat off-topic for the general Cygwin list...
I've been frustrated at the lack of guidance I've found online regarding
how to build an internationalized text-based Unix application, in particular
with reference to handling different character sets. For example, "man"
recognizes the character set specified in locale settings and uses Unicode
punctuation appropriately. OTOH, the cygutils "ascii -e" utility does not
recognize that my LANG specifies UTF-8 and outputs garbage for the
extended half. Should this be considered a bug?
Closer to a real application, how should a program that reads, processes,
and outputs text detect the character set so that it can process the text
correctly? In the absence of a BOM in the input file, should the program
blindly presume the locale-specified character set or analyze the file to
guess it? And what advice is there for guessing? And, once detected, what
takes precedence for output: the input file's character set or the locale
setting?
I presume these issues are either settled or being debated somewhere.
--Ken Nellis
next reply other threads:[~2011-05-23 13:44 UTC|newest]
Thread overview: 3+ messages / expand[flat|nested] mbox.gz Atom feed top
2011-05-23 13:44 Nellis, Kenneth [this message]
2011-05-23 14:51 ` Charles Wilson
2011-05-24 2:32 ` Buchbinder, Barry (NIH/NIAID) [E]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=2BF01EB27B56CC478AD6E5A0A28931F202910C7C@A1DAL1SWPES19MB.ams.acs-inc.net \
--to=kenneth.nellis@acs-inc.com \
--cc=cygwin-talk@cygwin.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for read-only IMAP folder(s) and NNTP newsgroup(s).