{"id":24589,"date":"2011-09-01T00:00:00","date_gmt":"2011-09-01T00:00:00","guid":{"rendered":"https:\/\/alienroad.com\/google-bilgi-bankasi\/pdfs-in-google-search-results\/"},"modified":"2011-09-01T00:00:00","modified_gmt":"2011-09-01T00:00:00","slug":"pdfs-in-google-search-results","status":"publish","type":"ar_kb","link":"https:\/\/alienroad.com\/google-bilgi-bankasi\/pdfs-in-google-search-results\/","title":{"rendered":"PDFs in Google search results"},"content":{"rendered":"<p class=\"gargardate\">Thursday, September 01, 2011<\/p>\n<p>\n  Our mission is to organize the world&#8217;s information and make it universally accessible and useful.<br \/>\n  During this ambitious quest, we sometimes encounter non-HTML files such as PDFs, spreadsheets, and<br \/>\n  presentations. Our algorithms don&#8217;t let different filetypes slow them down; we work hard to<br \/>\n  extract the relevant content and to index it appropriately for our search results.  But how do we<br \/>\n  actually index these filetypes, and\u2014since they often differ so much from standard HTML &mdash;<br \/>\n  what guidelines apply to these files? What if a webmaster doesn&#8217;t want us to index them?\n<\/p>\n<p><img decoding=\"async\" alt src=\"https:\/\/alienroad.com\/wp-content\/uploads\/kb-gorsel\/g-04c9cc9dd072.png\" loading=\"lazy\"><\/p>\n<p>\n  Google<br \/>\n  <a href=\"https:\/\/searchenginewatch.com\/2163391\" class=\"external-link\">first started indexing PDF files in 2001<\/a><br \/>\n  and currently has<br \/>\n  <a href=\"https:\/\/www.google.com\/search?q=filetype:pdf\" class=\"external-link\">hundreds of millions of PDF files indexed<\/a>.<br \/>\n  We&#8217;ve collected the most often-asked questions about PDF indexing; here are the answers:\n<\/p>\n<p>\n  <b>Q: Can Google index any type of PDF file?<\/b><br \/>\n  A: Generally we can index textual content (written in any language) from PDF files that use<br \/>\n  various kinds of character encodings, provided they&#8217;re not password protected or encrypted. If the<br \/>\n  text is embedded as images, we may process the images with<br \/>\n  <a href=\"https:\/\/googleblog.blogspot.com\/2008\/10\/picture-of-thousand-words.html\" class=\"external-link\">OCR<\/a><br \/>\n  algorithms to extract the text. The general rule of the thumb is that if you can copy and paste<br \/>\n  the text from a PDF document into a standard text document, we should be able to index that text.\n<\/p>\n<p>\n  <b>Q: What happens with the images in PDF files?<\/b><br \/>\n  A: Currently the images are not indexed. In order for us to index your images, you should create<br \/>\n  HTML pages for them. To increase the likelihood of us returning your images in our search results,<br \/>\n  please read the<br \/>\n  <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/images-best-practices\/\">Google Images best practices<\/a>.\n<\/p>\n<p>\n  <b>Q: How are links treated in PDF documents?<\/b><br \/>\n  A: Generally links in PDF files are treated similarly to links in HTML: they can pass PageRank<br \/>\n  and other indexing signals, and we may follow them after we have crawled the PDF file. It&#8217;s<br \/>\n  currently not possible to use<br \/>\n  <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/rel-attributes\/\"><code>nofollow<\/code><\/a><br \/>\n  links within a PDF document.\n<\/p>\n<p>\n  <b>Q: How can I prevent my PDF files from appearing in search results; or if they already do, how<br \/>\n    can I remove them?<\/b><br \/>\n  A: The simplest way to prevent PDF documents from appearing in search results is to add an<br \/>\n  <code>X-Robots-Tag: noindex<\/code> in the HTTP header used to serve the file. If they&#8217;re already<br \/>\n  indexed, they&#8217;ll drop out over time if you use the <code>X-Robot-Tag<\/code> with the<br \/>\n  <code>noindex<\/code> rule. For faster removals, you can use the<br \/>\n  <a href=\"https:\/\/www.google.com\/support\/webmasters\/bin\/answer.py?answer=164734\" class=\"external-link\">URL removal tool<\/a><br \/>\n  in Google Webmaster Tools.\n<\/p>\n<p>\n  <b>Q: Can PDF files rank highly in the search results?<\/b><br \/>\n  A: Sure! They&#8217;ll generally rank similarly to other webpages. For example, at the time of this<br \/>\n  post,<br \/>\n  <a href=\"https:\/\/www.google.com\/search?q=mortgage%20market%20review\" class=\"external-link\">mortgage market review<\/a>,<br \/>\n  <a href=\"https:\/\/www.google.com\/search?q=irs%20form%202011\" class=\"external-link\">irs form 2011<\/a>, or<br \/>\n  <a href=\"https:\/\/www.google.com\/search?q=paracetamol%20expert%20report\" class=\"external-link\">paracetamol expert report<\/a><br \/>\n  all return PDF documents that manage to rank highly in our search results, thanks to their content<br \/>\n  and the way they&#8217;re embedded and linked from other webpages.\n<\/p>\n<p>\n  <b>Q: Is it considered duplicate content if I have a copy of my pages in both HTML and PDF?<\/b><br \/>\n  <br \/>\n  A: Whenever possible, we recommend serving a single copy of your content. If this isn&#8217;t possible,<br \/>\n  make sure you indicate your preferred version by, for example, including the preferred URL in your<br \/>\n  sitemap or by specifying the canonical version in the HTML or in the<br \/>\n  <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/specify-canonical\/\">HTTP headers<\/a><br \/>\n  of the PDF resource. For more tips, read our Help Center article about<br \/>\n  <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/specify-canonical\/\">canonicalization<\/a>.\n<\/p>\n<p>\n  <b>Q: How can I influence the title shown in search results for my PDF document?<\/b><br \/>\n  A: We use two main elements to determine the title shown: the title metadata within the file, and<br \/>\n  the anchor text of links pointing to the PDF file. To give our algorithms a strong signal about<br \/>\n  the proper title to use, we recommend updating both.\n<\/p>\n<p>\n  If you want to learn more, watch Matt Cutt&#8217;s video about<br \/>\n  <a href=\"https:\/\/www.youtube.com\/watch?v=oDzq-94lcWQ\" class=\"external-link\">PDF files&#8217; optimization for search<\/a>,<br \/>\n  and visit our<br \/>\n  <a href=\"https:\/\/www.google.com\/support\/webmasters\/bin\/answer.py?answer=35287\" class=\"external-link\">Help Center<\/a><br \/>\n  for information about the content types we&#8217;re able to index. If you have feedback or suggestions,<br \/>\n  please let us know in the<br \/>\n  <a href=\"https:\/\/support.google.com\/webmasters\/community\" class=\"external-link\">Webmaster Help Forum<\/a>.\n<\/p>\n<p class=\"byline-author\">\n  Posted by<br \/>\n  <a href=\"https:\/\/garyillyes.com\/+\" rel=\"author\" class=\"external-link\">Gary Illyes<\/a>,<br \/>\n  Webmaster Trends Analyst<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Thursday, September 01, 2011 Our mission is to organize the world&#8217;s information and make it universally accessible and useful. During this ambitious quest, we sometimes encounter non-HTML files such as PDFs, spreadsheets, and presentations. Our algorithms don&#8217;t let different filetypes slow them down; we work hard to extract the relevant content and to index it [&hellip;]<\/p>\n","protected":false},"menu_order":84782,"template":"","meta":{"footnotes":""},"ar_kb_kategori":[665],"ar_kb_etiket":[],"class_list":["post-24589","ar_kb","type-ar_kb","status-publish","has-post-thumbnail","hentry","ar_kb_kategori-blog"],"_links":{"self":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb\/24589","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb"}],"about":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/types\/ar_kb"}],"version-history":[{"count":0,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb\/24589\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/media\/26898"}],"wp:attachment":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/media?parent=24589"}],"wp:term":[{"taxonomy":"ar_kb_kategori","embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb_kategori?post=24589"},{"taxonomy":"ar_kb_etiket","embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb_etiket?post=24589"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}